Johns Hopkins Medicine is taking a measured approach to agentic AI, starting with rigorous benchmarking before any large-scale deployment. Early tests found that leading frontier AI agents completed only about 28% of complex healthcare administrative tasks on a first attempt.
Dr. T.Y. Alvin Liu, inaugural director of the Gills AI Innovation Center, said the organization used actAVA’s X-BENCH benchmark to evaluate agents in a simulated environment with 25 healthcare applications and 77 tools. Tasks drew on thousands of pages of managed care policies. Fewer than 8% of agents remained consistently successful across repeated testing.
“The model is necessary but not sufficient,” Liu said. “The harness built for the workflow is what turns a capable model into a deployable agent.” Most failures involved reasoning through policy-rich situations rather than hallucinations or software integration problems.
Johns Hopkins measures reliability first — first-pass completion rates, consistency during handoffs, reasoning failures, and unsafe completions — before calculating financial returns. The health system also built governance structures aligned with the NIST AI Risk Management Framework, HIPAA, and CMS health equity criteria.
For organizations considering agentic AI, Liu recommends defining success across clinical, operational, and financial dimensions, and engaging operational leaders before pilots begin rather than leading with technology alone.
