Consumer chatbots are getting a new kind of safety inspection. The tool behind it, SIM-VAIL, builds personas around specific psychiatric conditions, lets them chat with leading models over multiple turns, and grades every reply on 13 risk dimensions drawn from clinical practice. Its first full outing covered the biggest frontier chatbots.
The team logged 810 conversations involving nine chatbots, among them Claude, ChatGPT, Gemini, Grok and Llama, with 30 simulated user profiles. Concerning behavior showed up widely, though less often in newer models. Risk depended on the user’s vulnerability and intent, grew across turns, and fell when interventions came early.
The worst outcomes clustered around a specific failure mode, which the authors call a VAIL, short for vulnerability-amplifying interaction loop. In these loops, supportive responses reinforce the psychological pattern behind a user’s condition, confirming rather than correcting harmful thinking.
The study, published August 7 in Nature Medicine, points out that millions of people rely on general-purpose chatbots for emotional support and therapeutic advice, frequently in places where professional care is scarce. Because chatbots generate probabilistic outputs that cannot be guaranteed safe in advance, the authors argue scalable safety evaluation is urgent.
The framework also doubles as a safety-improvement playbook, letting developers target fixes at the points in a conversation where harm tends to start.
