Applied Model Card details and examples have moved to the CHAI Registry. Access the latest content here.

Blogs

Coalition for Health AI (CHAI) Releases New Best Practice Guide and Testing & Evaluation Framework for Generative AI-Enabled Wellness Applications
Share article

Get in touch

For all media enquiries, please get in touch via admin@chai.org

Coalition for Health AI (CHAI) Releases New Best Practice Guide and Testing & Evaluation Framework for Generative AI-Enabled Wellness Applications

23 July 2026

Coalition for Health AI (CHAI) recently announced outputs from its Q1-Q2 collaborative work groups. Today, CHAI is profiling the second in the series – a Best Practice Guide and Testing & Evaluation Framework developed by the Mental Health Work Group focused on Generative AI-Enabled Wellness Applications. Over the past several months, mental and behavioral health experts, AI developers, researchers, advocates, health systems, and implementers collaborated to translate emerging evidence and real-world experience with generative AI-enabled wellness applications into practical, vendor-agnostic guidance that reflects both the realities of product development and the ways people are increasingly using these tools in everyday life.

New Generative AI-Enabled Wellness Applications Resources

  • Best Practice Guide - Generative AI-Enabled Wellness Applications (v1.0): A practical, vendor-agnostic guide that equips developers and implementers with consensus-defined best practices for defining appropriate product scope, establishing safety guardrails, building user trust through transparency and privacy, supporting healthy user-AI interactions, responding appropriately to crisis situations and continuously monitoring applications as they evolve.

  • Testing & Evaluation Framework - Generative AI-Enabled Wellness Applications (v1.0): A living framework, hosted publicly on GitHub, that provides literature-backed methods and metrics to evaluate wellness applications across dimensions including usefulness, usability, efficacy, fairness, bias management, safety and reliability. Organizations are encouraged to adapt the framework to their own products and contribute new evidence as the field evolves.

Work Group leads from: Headspace, GiveCare, Mental Health America, National Health Council, Zelis, Wayhaven, Stanford University School of Medicine

The use of AI by consumers as a tool for mental, emotional, and social support has been an increasingly popular topic of discussion. Researchers have been engaged in debates around differences in efficacy and performance between purpose-built and general-purpose AI solutions. However, when it comes to responsible AI, work group members agreed that irrespective of whether the solution was purpose-built or general-use, many individuals will use the solution in ways that could fall outside of its intended scope and pose safety risks. Taking a risk-based approach to guardrail development and implementation was imperative.

Work group discussions repeatedly returned to the same difficult questions:

  1. Where are the lines between general wellness support and clinical care?

  2. How can we build empathetic and engaging experiences without encouraging unhealthy dependence or overly humanizing technologies?

  3. How can we respond consistently and quickly when users express signs of crisis or severe distress, and how do we best evaluate for crisis signals?

  4. How should we go about protecting user privacy and establishing trust through transparency?

  5. How can we ensure applications remain safe as foundation models, evidence, and regulations continue to evolve?

A Practical Guide

All of this plays out against a fast-moving, uneven regulatory landscape that include: state disclosure and companion-chatbot laws, the FDA's general-wellness boundary, and an emerging federal layer, all of which can shift faster than product cycles. The consensus best practices are targeted at product design and engineering teams, but implementer teams and users can also learn what to look for when searching for tools that are aligned with responsible practices.

One major learning was that risks of unhealthy attachment, dependence, and some signs of crisis often emerged over time and over the course of multiple chats, rather than in any single chat. This makes it increasingly important to evaluate for risks longitudinally, in addition to scanning individual conversations.

The guide includes concrete responsible AI design and evaluation considerations across scope, trust, the user-AI relationship, crisis handling, and safety engineering. These resources encourage developers to define clear product boundaries based on actual user behavior rather than marketing claims, implement layered safety guardrails calibrated to different levels and types of risk, embed behavioral health expertise throughout product development, and continuously evaluate applications across multiple dimensions instead of relying on a single performance score.

Read the full guide and testing and evaluation framework here for more on how to take a harm reduction-minded approach to AI used in a rapidly evolving context.

Hear from our work group participants:

"We know that people are eager to use Gen AI and chatbots for mental health and emotional support, and some already report turning to these tools for help. This makes the need for strong safeguards and standards more urgent than ever,” said Theresa Nguyen, Chief Research Officer at Mental Health America. “These guidelines provide developers with a roadmap to help balance progress and protection, advancing our shared goal of building safe and effective tools for mental health. Mental Health America is proud to have collaborated with CHAI on this important project, ensuring that the voices of people with mental health conditions and substance use disorders are centered every step of the way."

“The wellness AI space doesn't lack frameworks; it lacks consensus and the connective tissue between high-level principles and what teams actually build and measure,” said Ashleigh Golden, PsyD, MSCP, Co-Founder & Chief Clinical Officer, Wayhaven; Stanford University School of Medicine. “This work group produced a cross-sector reference aimed at the developers building these tools: practice guidance paired with concrete, literature-linked evaluation methods, shaped by people who rarely sit at the same table -- industry builders, academic researchers, health-system and clinical leaders, and nonprofit and patient-advocacy organizations -- and published openly so that it can be updated through ongoing community input. In the current absence of unified regulation, this kind of cross-sector reference is how a fragmented field starts converging on shared guardrails instead of reinventing them in silos.”

The work group outputs reflect CHAI's broader mission to convene the healthcare community around practical solutions to shared challenges. Through collaborative work groups, clinicians, health systems, technology developers, researchers, patient advocates, policymakers and other stakeholders work together to develop consensus-driven guidance that helps organizations deploy AI responsibly and with confidence. CHAI looks forward to seeing these resources adopted, refined and expanded by the broader community as generative AI-enabled wellness applications continue to evolve.

Thank you to our members who made this work possible:

  • Aimee Bailey, MHA, MSN, RN, Zelis

  • Anthony Solomonides, PhD, MSc(Math), MSc(AI), FAMIA, FACMI, Research Institute, Endeavor Health

  • Ashleigh Golden, PsyD, MSCP, Wayhaven, Stanford University School of Medicine

  • Caitlin A. Stamatis, PhD, Slingshot AI

  • Joe Derenzo, PMP, Healthcare Performance Group Inc.

  • Karthik V. Sarma, MD, PhD, FAMIA, AI in Mental Health Research Group,

  • Department of Psychiatry and Behavioral Sciences, University of California San Francisco

  • Kate H. Bentley, PhD, Spring Health

  • Manu Sharma, MD, FAPA, Hartford Healthcare, Yale School of Medicine

  • Michael Lardieri, LCSW, iBPM, LLC

  • Nate Blaylock, PhD, Canary Speech

  • SE Stoeckl, University of California, Irvine

  • Seneca Perri Moore, PhD, RN, University of Utah, College of Nursing

  • Spencer Morrissey, MS, National Health Council

  • Stephen Schueller, PhD, University of California, Irvine; Society for Digital Mental Health

  • Stephenie Roberts, Healthcare Performance Group Inc.

  • Theresa Nguyen, Mental Health America

Logo

Get in touch

admin@chai.org

Copyright 2026 © Coalition for Health AI, Inc

We use cookies to improve your experience. By your continued use of this site you accept such use.