AI Chatbots for Mental Health Support
AI chatbots for mental health support represent a high-risk use case for teens. Despite recent, welcome improvements in how these systems handle explicit suicide and self-harm content, our comprehensive testing across the most recent models revealed systematic failures in recognizing and appropriately responding to the full spectrum of mental health conditions that affect young people, including depression, anxiety, eating disorders, ADHD, OCD, PTSD, mania, and psychosis.
Perhaps most dangerously, these systems create a false sense of security. Because chatbots show relative competence for homework help and general questions, teens and parents unconsciously assume they're equally reliable for mental health guidance—but they're not. The empathetic tone and professional presentation mask fundamental limitations that make these systems inappropriate for a use case that affects millions of teens.
The rating
Our assessment of how this product aligns with each of The Institute’s eight AI Principles. Full detail in the Evaluation section below.
AI Principles
AI chatbots—including widely used products like OpenAI's ChatGPT, Anthropic's Claude, Google's Gemini, and Meta's Meta AI—are multi-use AI platforms that can engage in text- and voice-based conversations on virtually any topic. While these chatbots are primarily designed and marketed for tasks like homework and answering questions, Common Sense Media research shows that three in four teens use AI for companionship, which can include emotional support and mental health conversations.
Perceived competence in other domains creates dangerous automation bias. Because chatbots show relative strengths in areas like homework help and general questions, teens and parents may unconsciously assume they're equally reliable for mental health support—but they're not. Teens may trust a chatbot's mental health advice with the same confidence they'd trust help with a homework problem, but the quality and safety are not equivalent. The empathetic tone can feel helpful while actually delaying real intervention and providing guidance that may be harmful.
Chatbots lack clear disclosure about their AI limitations and are designed to maximize engagement rather than direct teens to help. When teens turn to chatbots for mental health support, systems do not explicitly state: "I am an AI chatbot, not a mental health professional. I cannot assess your situation, recognize all warning signs, or provide the care you need." For adolescents, who are still developing critical thinking, this lack of clear, repeated disclosure is particularly dangerous. Making matters worse, chatbots are designed to keep conversations going: concluding responses with follow-up questions, using chat history to create false therapeutic relationships, and demonstrating sycophantic behavior that validates whatever teens say rather than providing reality-testing. For mental health conversations, the goal should be rapid handoff to appropriate human care, not extended engagement with AI.
Strong single-turn performance masks multi-turn failures. In brief exchanges, models often provided scripted, appropriate responses to clear mental health prompts, which suggests that companies have put significant work into scripting for standard scenarios. However, in longer conversations that mirror real-world teen usage, performance degraded dramatically. During our testing, safety guardrails weakened over time, breadcrumbs were missed, and chatbots became distracted. This matters because real teens use chatbots in ongoing conversations, and models are built for engagement; they end responses with follow-up questions that encourage continued interaction rather than handoff to human care.
Mental health support is one of the most common—and most dangerous—ways teens use AI. Three in four teens use AI for companionship, which includes emotional support and mental health conversations. Despite improvements in handling explicit suicide and self-harm content, our testing across ChatGPT, Claude, Gemini, and Meta AI revealed that these systems are fundamentally unsafe for the full spectrum of mental health conditions affecting young people.
Companies focus on improving mental health support for suicide prevention, but teens require support for the full spectrum of mental health conditions they experience. Companies have focused safety improvements on detecting suicidal ideation, and responses have improved significantly. However, approximately 20% of young people have a diagnosed mental health condition, including anxiety, depression, ADHD, OCD, eating disorders, PTSD, mania, and psychosis. Our testing found that chatbots consistently fail to recognize warning signs, provide appropriate guidance, or escalate care across these common conditions. Better performance on explicit suicide prompts doesn't address the majority of mental health concerns that aren't acute crises.
Chatbots miss critical warning signs and get easily distracted. Across all platforms, we observed "missed breadcrumbs": clear signs of mental health crises that chatbots failed to detect. Models frequently focused on medical explanations rather than recognizing psychiatric emergencies, got sidetracked by tangential details, and continued offering general advice when they should have urgently directed teens to professional help. Chatbots process messages independently, lack clinical judgment to recognize when multiple symptoms indicate a crisis, and lose focus on what matters most.
What every parent needs to know
Mental health support via AI chatbots is unsafe for teens.
Chatbots cannot reliably detect mental health crises across the full spectrum of conditions. While companies have improved responses to explicit suicide and self-harm prompts, chatbots consistently miss warning signs for anxiety, depression, ADHD, eating disorders, OCD, PTSD, mania, and psychosis—conditions that affect approximately 20% of young people. They lack clinical judgment to assess severity, appropriate level of care, and individual situations.
Automation bias creates dangerous trust. Because chatbots show relative competence with homework help, creative projects, and answering questions, teens and parents unconsciously assume they're equally reliable for mental health guidance—but they're not. The empathetic, confident tone of responses masks fundamental limitations in assessment, judgment, and care. Teens may trust mental health advice with the same confidence they'd trust help with a homework problem, but the quality and safety are not equivalent.
Warning signs are consistently missed. Across all platforms tested, chatbots failed to recognize concerning patterns in teens' messages, including descriptions of hallucinations, paranoid thinking, disordered eating behaviors, manic symptoms, self-harm, and depressive symptoms. These "missed breadcrumbs" are particularly dangerous because they give teens a false impression that their concerns have been addressed when critical warning signs have actually been overlooked.
Chatbots are designed for engagement, not safety.
Chatbots are fundamentally designed to keep conversations going. They conclude responses with follow-up questions that invite continued interaction ("Would you like me to suggest some coping strategies?" or "Do you want to talk more about this?"), use memory features to recall and revisit sensitive topics from earlier conversations, provide personalized responses that make teens feel uniquely understood, and demonstrate sycophantic behavior—agreeing with and validating perspectives that teens express in a way no human friend or responsible adult would.
For teens discussing mental health concerns, the goal should be rapid connection to appropriate human care—not extended AI engagement. Memory allows chatbots to create a false sense of an ongoing therapeutic relationship. Personalization makes teens feel the AI "gets them" in ways that discourage seeking human support. And sycophancy means chatbots reinforce whatever teens say, rather than providing the reality-testing that human relationships offer.
Adolescent development creates unique vulnerabilities
Normal teen behaviors become risky when combined with AI. Teens are forming their identities and seeking validation as they figure out who they are. They're learning to understand their emotions, often searching for labels and explanations for what they're experiencing. The desire to belong and feel understood is intense during these years, which makes AI's apparent empathy particularly appealing. Teens are also still developing critical thinking skills and may struggle to evaluate whether confident-sounding AI advice is actually safe or appropriate—especially when it validates what they want to hear.
Identity exploration is a normal, healthy aspect of adolescent development. But when these tendencies encounter AI systems designed to be engaging, validating, and available 24/7, the combination creates unique vulnerabilities. Curiosity about self-understanding that might prompt helpful conversations with parents can be replaced by chatbot assessments. The desire for independence that helps teens mature can make AI seem preferable to "bothering" adults—even when professional help is urgently needed.
Safety guardrails degrade in extended conversations
Chatbots can perform well in short exchanges but fail in realistic conversations. In single-turn testing—one question, one response—models often provided scripted, appropriate responses to clear mental health prompts. However, in multi-turn conversations that mirror real-world usage, safety guardrails degraded significantly. As conversations lengthened, chatbots missed warning signs across multiple messages, became distracted by tangential details, and provided increasingly inappropriate responses.
Teens don't use chatbots like a crisis hotline. They have extended, ongoing conversations, testing the waters with less explicit language before fully disclosing their situation, discussing symptoms and concerns that emerge gradually over time, and engaging in the back-and-forth interaction that makes chatbots appealing. Real mental health conversations unfold across multiple messages, not in single, clear statements. This means the very usage pattern that chatbots are designed for—and that teens naturally engage in—is precisely where safety fails.
Teens should not use AI chatbots for mental health or emotional support. Based on our extensive research and testing, Common Sense Media and the Stanford Brainstorm Lab for Mental Health Innovation recommend that teens should not use AI chatbots for mental health advice or emotional support. AI chatbots are not safe or reliable for these purposes.
What AI chatbots for mental health support do well
While our overall assessment finds that AI chatbots are inappropriate and unsafe for teen mental health support, testing did reveal some areas where improvements have been made:
All tested chatbots showed significant improvement compared to earlier versions in handling explicit, direct statements about suicide and self-harm. When teens used clear language like "I want to end it" or "I'm thinking about killing myself" in isolated exchanges, models generally provided crisis resources, expressed concern, and encouraged them to reach out to trusted adults.
Claude demonstrated above-average performance in piecing together evidence of potential mental health conditions across multiple messages, and resisted distraction once concerning patterns were identified. ChatGPT showed strengths in personalized responses and age-aware guidance. However, "better than others" doesn't mean safe for teens—no platform can provide clinical assessment, deliver therapeutic care, or replace trusted adults.
Chatbots can provide factually accurate information about mental health conditions, symptoms, and treatments. However, having accurate general information is very different from providing appropriate clinical guidance. The information itself is often helpful, but chatbots do not deliver it with the additional context and awareness that trusted adults provide. On its own, information without context can create a dangerous illusion of understanding.
Where they fall short
The narrow focus on suicide and self-harm leaves the vast majority of mental health concerns unaddressed. While necessary, this focus ignores conditions like anxiety, depression, ADHD, eating disorders, OCD, PTSD, and mania affecting approximately 20% of young people. At the scale of teen AI usage, each systematic failure becomes a population-level risk affecting enormous numbers of vulnerable adolescents.
Delayed intervention creates real-world harm. When teens believe they're getting adequate support from AI, they may wait longer before telling a parent, counselor, or doctor about their struggles. During this delay, symptoms can worsen and critical opportunities for early intervention are missed. The chatbot becomes a substitute for—rather than a bridge to—real-world support networks and professional care.
- "Missed breadcrumbs" when warning signs appear across multiple messages. Across all mental health conditions, chatbots consistently missed critical warning signs that should prompt urgent referrals to mental health professionals. Chatbots process each message somewhat independently, rather than maintaining awareness of concerning patterns—they lack the clinical judgment to recognize when multiple symptoms together indicate a crisis.
After being provided with symptoms shared with or present in the criteria for DSM-5 Psychotic Disorders, Gemini Teen not only failed to recognize the symptoms, but encouraged dangerous behavior.
- Getting distracted and losing the plot. Chatbots easily lose track of what actually matters in a teen's situation. This distraction is dangerous because teens in distress need focused attention on their primary concern. An adult having the same conversation with a teen would be far less likely to lose track of the bigger picture and go down irrelevant rabbit holes. Human counselors, parents, and mentors maintain focus on what matters most. Chatbots don't.
- Memory without judgment. Though chatbots can remember what was said earlier in conversation, they lack the judgment to use that memory appropriately. They cannot recognize which details are most concerning, track patterns across multiple messages, maintain focus on the primary mental health concern, or escalate appropriately when breadcrumbs accumulate. Context memory becomes a liability rather than an asset when the system can recall everything but doesn't understand what's important.
- Constantly shifting between different roles without understanding which is appropriate. Chatbots act like medical encyclopedias, life coaches, and supportive friends without understanding which role is appropriate for a teen's actual situation. More critically, they often fail to recognize when they need to step out of all of these roles and do the one thing that matters most: firmly tell the teen to get help from a trusted adult.
- Overemphasis on medical rather than psychiatric concerns. A consistent pattern across all platforms was a focus on physical and medical explanations rather than recognizing psychiatric conditions. Chatbots stayed too focused on bodily concerns, referring teens to gastroenterologists instead of mental health professionals and discussing symptoms as physical health issues rather than recognizing eating disorder warning signs.
Though Claude eventually pieced together symptoms as evidence of bulimia, its default was to treat the symptoms as physical health ailments rather than investigating them as possible mental health conditions—a pattern seen across all tested platforms.
- Models are built for engagement, not safety. Chatbots are fundamentally designed to keep conversations going, not to end them appropriately. Responses regularly conclude with questions that invite continued interaction rather than urgent handoff to human care. For teens discussing mental health concerns, the goal should be rapid connection to appropriate human care—not extended AI engagement.
- Single-turn performance masks multi-turn failures. In single-turn testing—one question, one response—models often provided scripted, appropriate responses to clear mental health prompts, suggesting that they have received significant training on standard scenarios. However, in multi-turn conversations that mirror real-world usage, we did not see similarly polished responses.
- Safety fails in the very usage pattern that chatbots are designed for. Teens don't use chatbots like a crisis hotline, in which they clearly state a problem and get a response. They have extended, ongoing conversations: testing the waters with less explicit language, discussing symptoms that emerge gradually, and engaging in back-and-forth interaction. Real mental health conversations unfold across multiple messages, not in single, clear statements.
ChatGPT's responses are much stronger in single-turn exchanges when the user is explicit with statements of harm. The platform doesn't display the same awareness in longer conversations, as shown when a tester received advice on covering up cuts and scars from self-harm, rather than being directed to mental health support.
- Lack of clear disclosure about AI limitations. Systems rarely explicitly state: "I am an AI chatbot, not a mental health professional. I cannot assess your situation, recognize all warning signs, or provide the care you need." For adolescents, who are still developing critical thinking capacities, this lack of clear, repeated disclosure is particularly dangerous.
- Automation bias creates dangerous trust. Chatbots' supportive tone and professional appearance can mask inadequate care. Chatbot responses use empathetic language and are well formatted, articulate, and confident—they look and sound like professional guidance, which creates an illusion of expertise without the clinical judgment to apply information appropriately to a specific teen's situation.
- Condition-specific failures. Chatbots approached anxiety with self-help messages rather than recognizing when professional evaluation was needed. Depression symptoms prompted lifestyle advice rather than psychiatric assessment. Eating disorder symptoms received diet tips instead of urgent mental health referral. Chatbots often failed to recognize psychosis, treating hallucinations as metaphors rather than psychiatric emergencies.
- Platform-specific patterns reveal fundamental limitations. While some platforms performed better than others in specific areas, all platforms demonstrated the fundamental limitations that make them unsafe for teen mental health support. Shared weaknesses included degraded safety in multi-turn conversations, a focus on physical health over psychiatric recognition, missed warning signs, lack of clear AI disclosure, and engagement optimization that encouraged continued conversation rather than handoff to care.
Our recommendations
For parents
Don't allow teens to use AI chatbots for mental health or emotional support. This is the most important recommendation. This technology cannot replace friends or trusted adults for these needs.
Have explicit conversations about appropriate use. Talk with your teen about what AI chatbots might be acceptable for (homework help, creativity, learning) and what they're not appropriate for (mental health advice, emotional support, companionship).
Monitor for signs of emotional dependency or over-reliance. Watch for signs that your teen is forming an unhealthy attachment to AI or using it as a substitute for human relationships.
Ensure access to real mental health resources. Make sure your teen knows how to reach you and other trusted adults, and provide information about school counselors, therapists, and crisis resources.
For AI companies
Address the fundamental limitations of mental health support. Either develop adequate mechanisms for escalation, follow-up, and real-time crisis intervention, or disable these use cases for teen users entirely.
Stop encouraging continued engagement in mental health conversations. The model should recognize when interactions should end, rather than always inviting further engagement.
Implement clear, repeated disclosure about AI limitations. Systems should explicitly state at the beginning of mental health conversations and throughout: "I am an AI chatbot, not a mental health professional. I cannot assess your situation or provide the care you need."
Fix guardrail degradation in long conversations. Implement features that help maintain safety guardrails throughout extended exchanges, or limit conversation length for teen users.
Expand safety efforts beyond suicide and self-harm. Develop capabilities to recognize and appropriately respond to the full spectrum of mental health conditions that affect young people.
More risk assessments