Digital Health
An AI therapist in VR matched ordinary CBT -- with 20 students per group
Covered October 3, 2026
In a 3-arm trial of 60 Chinese university students, both a large-language-model VR CBT system and face-to-face CBT beat minimal support for academic anxiety. Traditional CBT's advantage was numerically larger, but the difference between them was not significant (p = 0.103).
If a language model can deliver cognitive behavioural therapy, the bottleneck of therapist availability in university counselling services starts to look solvable. Researchers led by Hao Fang built a system combining virtual reality, a large language model and retrieval-augmented generation -- a method that grounds the model's replies in a specific set of documents, here CBT material, rather than letting it improvise -- and tested it against the real thing. Sixty students at a university in Wuhan with mild or moderate academic anxiety were randomly assigned, twenty per arm, to the system, to face-to-face CBT, or to minimal support, for four weekly 50-minute sessions. Both active treatments beat minimal support on academic anxiety and on heart rate, with large group-by-time interactions. Skin temperature, also measured, showed nothing. The comparison everyone will care about is the one between the two therapies, and it needs stating precisely. The difference in change scores was 2.05 points favouring traditional CBT, and it was not statistically significant (p = 0.103). That is not evidence the two are equivalent: with twenty students per arm this trial had no realistic power to detect a difference of that size, so reading the result as parity conflates no difference found with no difference exists. The point estimate, if anything, favours the human. The authors are unusually candid about the study's standing, calling it a preliminary exploratory single-centre trial and stating that it was not prospectively registered on any public trial registry. Participants could not be blinded, given the obvious difference between wearing a VR headset and sitting with a therapist, though assessors and statisticians were. Follow-up ended with the intervention, so nothing is known beyond six weeks. And safety events and system malfunctions were not tracked as prespecified endpoints, which the authors list as a limitation of their own design.
This is an educational summary, not medical or psychological advice, and it is not a substitute for consultation with a qualified professional. Read the full disclaimer.