By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
MedsparkMedsparkMedspark
  • Home
  • News & Alerts
    News & AlertsShow More
    Google SymptomAI outperforms doctors in real-world symptom interviews
    By
    msadmin
    July 27, 2026
    White House convenes clinical AI experts for standards sprint on evaluation
    By
    msadmin
    July 27, 2026
    AI shrinks drug discovery timelines with data-driven biologics design
    By
    msadmin
    July 27, 2026
    OpenAI launches Health in ChatGPT day after lawsuit targets its safety
    By
    msadmin
    July 27, 2026
    AI-focused venture firm Dimension restocks with $800M third fund
    By
    msadmin
    July 23, 2026
  • Spotlight
    SpotlightShow More
    OpenAI launches GPT-Rosalind for life sciences research
    By
    msadmin
    July 17, 2026
    ARPA-H Awards $160 Million for AI-Enabled Personalized Gene Editing to Tackle Rare Diseases
    By
    msadmin
    July 12, 2026
    Strategic healthcare AI governance framework abstract illustration
    Building a Resilient Healthcare AI Strategy: Insights from Industry Leaders
    By
    msadmin
    May 15, 2026
    Pharma AI Alliance Expands: Owkin and AstraZeneca Deploy New Drug Discovery Models
    By
    msadmin
    May 14, 2026
    7 Must-Attend MedTech Events in South Africa for 2025
    By
    Jostel Owusu
    August 9, 2025
  • Articles
    ArticlesShow More
    Teladoc puts 100 percent of fees at risk with new AI-powered platform
    By
    msadmin
    July 27, 2026
    Johns Hopkins finds frontier AI agents fail most complex health tasks
    By
    msadmin
    July 27, 2026
    Teladoc puts fees at risk in AI-powered virtual care platform relaunch
    By
    msadmin
    July 27, 2026
    AI tops global healthcare IT investment priorities, KLAS finds
    By
    msadmin
    July 22, 2026
    Hospital CEOs must own AI safety, not delegate it to IT
    By
    msadmin
    July 22, 2026
  • Events
    EventsShow More
    Stanford Health AI Week Highlights AI’s Growing Role in Medical Education, Patient Empowerment, and Life Sciences
    By
    msadmin
    June 19, 2026
    HIMSS APAC 2026: Re-engineering APAC Health Systems in the AI Era
    By
    msadmin
    June 9, 2026
    AIMed 2026: Bridging the Gap Between AI Promise and Clinical Reality in Kraków
    By
    msadmin
    April 30, 2026
    7 Must-Attend MedTech Events in South Africa for 2025
    By
    Jostel Owusu
    August 9, 2025
    Cleveland Clinic’s First AI Summit Signals Bold Future for Healthcare
    By
    msadmin
    July 19, 2025
  • About
    • Mission
    • Services
    • Contact
Font ResizerAa
MedsparkMedspark
Font ResizerAa
  • Home
  • News & Alerts
  • Spotlight
  • Articles
  • Events
  • About
  • Quick Links
    • Home
    • News & Alerts
    • Spotlight
    • Articles
    • Events
  • About MedSpark
    • Our Purpose & Vision
    • Services
    • Contact
Follow US
Articles

Johns Hopkins finds frontier AI agents fail most complex health tasks

A benchmark test found leading AI agents completed only 28% of healthcare administrative tasks on a first attempt.

MedSpark Staff
By
msadmin
MedSpark Staff
Bymsadmin
Medical, Healthcare, & Biotech/Pharma AI News
Follow:
Published: July 27, 2026
Share
2 Min Read
SHARE

Johns Hopkins Medicine is taking a measured approach to agentic AI, starting with rigorous benchmarking before any large-scale deployment. Early tests found that leading frontier AI agents completed only about 28% of complex healthcare administrative tasks on a first attempt.

Dr. T.Y. Alvin Liu, inaugural director of the Gills AI Innovation Center, said the organization used actAVA’s X-BENCH benchmark to evaluate agents in a simulated environment with 25 healthcare applications and 77 tools. Tasks drew on thousands of pages of managed care policies. Fewer than 8% of agents remained consistently successful across repeated testing.

“The model is necessary but not sufficient,” Liu said. “The harness built for the workflow is what turns a capable model into a deployable agent.” Most failures involved reasoning through policy-rich situations rather than hallucinations or software integration problems.

Johns Hopkins measures reliability first — first-pass completion rates, consistency during handoffs, reasoning failures, and unsafe completions — before calculating financial returns. The health system also built governance structures aligned with the NIST AI Risk Management Framework, HIPAA, and CMS health equity criteria.

For organizations considering agentic AI, Liu recommends defining success across clinical, operational, and financial dimensions, and engaging operational leaders before pilots begin rather than leading with technology alone.

TAGGED:Agentic AIAI AgentsAI benchmarkingAI Governancehealth ITHealthcare AIJohns Hopkins
SOURCES:Healthcare IT News
Share This Article
Facebook Copy Link Print
MedSpark Staff
Bymsadmin
Follow:
Medical, Healthcare, & Biotech/Pharma AI News
banner-medspark-horiz

You Might Also Like

ArticlesSpotlight

EY Expert Urges Healthcare Leaders to Double Down on AI Amid Economic Uncertainty

By
Yu Chi Huang
July 11, 2025
Articles

NHS Unveils Radical 10-Year AI-Driven Plan to Transform Healthcare Delivery

By
Yu Chi Huang
July 8, 2025
Articles

Google SensorFM trains on one trillion minutes of wearable health data

By
msadmin
July 19, 2026
Events

AIMed 2026: Bridging the Gap Between AI Promise and Clinical Reality in Kraków

By
msadmin
April 30, 2026

AI news, analysis, and insights for healthcare, biotech, and pharma.

Facebook Twitter Youtube Linkedin
Quick Links
  • News & Alerts
  • Articles
  • Spotlight
  • Events
About Medspark
  • Mission
  • Services
  • Contact

© Copyright 2026 MedSpark. All rights reserved.

Privacy Policy | Legal