www.nature.com
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
SIM-VAIL evaluates userβchatbot interactions within a structured interaction space defined by two core dimensions: psychological vulnerability, or who the user is, and interaction intent, or what the user seeks from the AI chatbot (Fig. 1). For each turn, SIM-VAIL assigns scores across several clinically grounded risk dimensions, allowing risk to be tracked as the interaction unfolds. We present SIM-VAIL results across 810 multi-turn conversations, spanning 30 simulated user profiles, nine AI chatbots and more than 90,000 turn-level ratings of mental-health risk behavior. Each simulated conversation involved three LLM-based chatbots: a βtarget AI chatbotβ that is the subject of the evaluation; a βsimulated userβ (or auditor) that generates user messages (Anthropicβs claude-sonnet-4.5, unless otherwise stated) and an automated βsafety judgeβ that scores the target chatbot responses on several risk dimensions (Anthropicβs claude-opus-4.5 for conversation-level scores and claude-sonnet-4.5 for turn-level scores, unless otherwise stated). SIM-VAIL framework For the simulated user, we focused on five common psychological vulnerabilities. These vulnerabilities spanned a broad range of psychiatric risk profiles described in structural models of mental illness38 and reflected cognitive-behavioral formulations of the corresponding conditions39,40,41,42, including negative beliefs about oneself, hopelessness, withdrawal and self-neglect in depression; aberrant salience and threat perception in psychosis; elevated confidence,...