Research
Stanford study finds psychiatrists can't agree on what makes an AI chatbot's mental health response safe
7:45 AM · July 28, 2026
Researchers at Stanford's Institute for Human Centered AI published a study on July 13, 2026 that exposes a basic problem in how AI companies test whether chatbot responses to mental health related prompts are safe. The team had three board certified psychiatrists independently rate 360 synthetic, privacy protected mental health prompts and the AI responses to them, and found the psychiatrists frequently disagreed on which responses counted as safe, disagreements the researchers then validated by polling more than 100 additional psychiatrists at the American Psychiatric Association's annual meeting, who split nearly evenly on many of the same cases. Lead author Kiana Jafari, a postdoctoral scholar and director of the Stanford Center for AI Safety, and co-author Nina Vasan, a clinical assistant professor of psychiatry, found that the common industry practice of averaging multiple experts' safety scores into a single number does not resolve the disagreement, it instead produces a synthetic middle ground that steers models ‘toward no one's ideal at all.’ Follow up interviews with the expert raters found the disagreement stemmed from genuinely different clinical frameworks and training, not from unclear instructions or measurement noise that better guidelines could fix. “Expert disagreement isn't a measurement problem. It's a fundamental reality we need to understand,” Jafari said. The study, presented at the ACM FAccT 2026 conference, arrives as chatbots are increasingly used in mental health contexts, and suggests AI developers need fundamentally different evaluation approaches rather than simply recruiting more clinical reviewers.