Back to Research papers
Research paper index

One Year Later...The Harms Persist, But So Do We!

Annika Marie Schoene, Cansu Canca, Gautham Vijay Kumar, Anson Antony

arXiv:2606.23884Published June 22, 2026Updated June 29, 20260 citations
  • cs.CL
  • cs.AI

Abstract

General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain inadequate and inconsistent across clinical conditions. This study evaluates eight proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100\%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into publicly available settings (e.g., schools, search engines, and consumer chatbots) are particularly concerning.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.