Coalition for Health AI (CHAI) Releases New Best Practice Guide and Testing & Evaluation Framework for Agentic AI
29 July 2026
Agentic technology is the current frontier of AI – shifting the paradigm from reactive, human-prompted assistance to autonomous problem-solving. Today, CHAI is profiling a Best Practice Guide and Testing & Evaluation Framework developed by the Work Group focused on Agentic AI, the third in a recently announced series of outputs from its Q1-Q2 collaborative work groups.
Over the past several months, health system leaders, AI developers, researchers and implementers collaborated to translate emerging experience with agentic AI into practical, vendor-agnostic guidance that reflects both the rapid pace of agentic AI innovation and the unique technical, operational and governance requirements of healthcare.
New Agentic AI Resources
Best Practice Guide – Agentic AI (v1.0): A practical, vendor-agnostic guide that equips developers and implementers with consensus-defined protocol considerations for designing, deploying and governing autonomous and semi-autonomous AI agents in healthcare. Rather than prescribing a single technical implementation, the guide identifies healthcare-specific capabilities that should extend emerging agent protocols to support safe, interoperable and responsible AI.
Testing & Evaluation Framework – Agentic AI (v1.0): A living framework, hosted publicly on GitHub, that provides literature-backed methods and metrics to evaluate agentic AI across dimensions including usefulness, usability, efficacy, fairness, safety, reliability and business impact. Organizations are encouraged to adapt the framework to their own environments and continue refining evaluation approaches as both the technology and evidence base mature.
Work Group leads from: Prompt Opinion, Practice Fusion, EBSCO, Dyna AI, Elsevier, Duke, PointClickCare, Innovaccer, Xsolis
Key Challenges & Insights
As healthcare organizations begin moving beyond conversational AI toward systems capable of reasoning, planning, and taking action across complex workflows, work group members quickly recognized that existing agentic technologies weren’t built with healthcare in mind. While emerging protocols such as MCP and A2A establish a foundation for agent communication, they do not capture the healthcare-specific context, governance requirements and accountability needed for safe deployment.
The work group's discussions repeatedly returned to the same difficult questions:
How should autonomous agents communicate patient context consistently across multi-agent workflows?
Where should organizations establish appropriate boundaries between agent autonomy and human oversight?
How can healthcare organizations evaluate increasingly autonomous systems when no accepted gold standard for agent evaluation exists?
What information should be captured to make agent reasoning transparent, traceable and auditable?
How can developers safely deploy agents as low-code tools accelerate adoption while governance processes continue to evolve?
Governance is also being outpaced, too: one health system found roughly three times as many agents built in approved low-code tools as use cases that had cleared formal review. And in voice scheduling, as one member put it, every phone call is an edge case.
Rather than prescribe one architecture, the guide translates those gaps into healthcare-specific extensions that ride on top of A2A and MCP. The extensions are framed around the problem to be solved, since the technical approaches will change faster than the principles. A few examples include:
Make data protections travel with the data, so purpose-of-use metadata and a consent-scope object propagate through every tool call and handoff, and a downstream analytics agent must decline work the patient never consented to.
Require authenticated human attestation before agent output is committed to the EHR, and clearly mark which data points the agent added or changed.
Derive safety metrics by asking clinicians how they'd supervise a staff member doing the task, while also considering what this agent could get wrong that a person almost certainly wouldn't, like silently proceeding when it should have paused to escalate.
In voice deployments, treat escalation on a clinical red flag as unconditional and exempt from optimization.
Notably, the guide also cautions that not every workflow should be an agent at all; many are better served by deterministic, validated "skills," and adding autonomy where it isn't needed just creates cost and agent clutter.
Read the full guide and testing and evaluation framework here to see the entire product of these collaborative work group conversations and meetings specific to agentic capabilities. As always, we welcome feedback and continued work with the community to build tools and resources to make guides practical and implementable.
Hear from our work group participants:
“As a practicing physician, it is critical to me that AI entering clinical workflows is safe, trustworthy, and held to a standard appropriate for healthcare,” said Kate Eisenberg, MD, PhD, Senior Medical Director of Dyna AI. “CHAI’s Best Practice Guide and Testing & Evaluation Framework for Agentic AI bring together real clinical examples and current evaluation methods to help define how AI agents should be built, assessed, and used responsibly. In a rapidly moving space, these resources give the field a shared foundation for advancing innovation while protecting patients and clinicians.”
"I think it's exhilarating to meet with other members of the industry and recognize through presentations, testimonies and conversations the patterns we all deal with on a daily basis,” said Olivier Rousseau, Technical Product Manager at Clinia. “It also feels refreshing to get non-academical views on the very real problems developers and implementers have, that only arise in production environments a lot of the time."
Thank you to our members who made this work possible:
Ashish Patel, MHS, CareSet
Pawan Jindal, MBBS, MS, Darena Health
Anand Chowdhury, MD, MMCi, Duke University School of Medicine
Katherine Eisenberg, MD, PhD, Dyna AI
Benjamin Hollis, EBSCO, Inc.
Daniel D. Johnson, BSN, RN, NI-BC, Essentia Health
Muhammad Xhemali. PharmD, RPh, Family Health Center of Worcester
Ashley M Hopkins, Flinders University
Sonali Tamhankar, PhD, Fred Hutchinson Cancer Center
Cheryl Campbell, Healthcare Performance Group Inc.
Joe Derenzo, PMP, Healthcare Performance Group Inc.
Michael Lardieri, LCSW, iBPM, LLC
Rémy Ibarcq, Individual Contributor
Dorcas Yao, MD, Individual Contributor
Michael Choma, MD, PhD, Infinitus Systems, Inc.
Tapan Shah, PhD, Innovaccer
Bita Behrouzi, MD, MaineHealth Maine Medical Center; Tufts University School of Medicine
Jeremy Attermann, MSW, National Council for Mental Wellbeing
Scott Ross, Opala
Shiba Kuanar, University of Minnesota
Seneca Perri Moore, PhD, RN, University of Utah, College of Nursing
Keri Hall, MD, MSc, Wolters Kluwer - UpToDate
