By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
MedsparkMedsparkMedspark
  • Home
  • News & Alerts
    News & AlertsShow More
    Medicare’s AI approval pilot drew warnings before it launched
    By
    msadmin
    September 16, 2026
    GE HealthCare’s new AI tools forecast hospital bottlenecks 72 hours out
    By
    msadmin
    September 16, 2026
    Knowtex bets on evaluation science to sell clinical AI to hospitals
    By
    msadmin
    September 16, 2026
    Pediatric heart models now take seconds instead of a workday
    By
    msadmin
    September 16, 2026
    OpenAI Foundation puts $40M behind cancer vaccine data collection
    By
    msadmin
    September 16, 2026
  • Spotlight
    SpotlightShow More
    Robinhood Ventures Fund II brings retail capital to healthcare AI
    By
    msadmin
    August 13, 2026
    Adialante brings accessible MRI-based cancer screening
    By
    msadmin
    August 11, 2026
    CellType models biology so AI can discover drugs
    By
    msadmin
    August 11, 2026
    Healthcare AI startups in Robinhood Ventures Fund II
    By
    msadmin
    August 11, 2026
    OpenAI launches GPT-Rosalind for life sciences research
    By
    msadmin
    July 17, 2026
  • Articles
    ArticlesShow More
    Proteomics model picks breast cancer drugs from biopsies
    By
    msadmin
    September 10, 2026
    CRISP model reads frozen sections to steer cancer surgery
    By
    msadmin
    September 10, 2026
    Consumer chatbots are building a medical system outside hospitals
    By
    msadmin
    August 21, 2026
    Reasoning gaps hold back AI agents in scientific discovery
    By
    msadmin
    August 11, 2026
    Benchmark scores can’t track real clinical LLM use, Stanford says
    By
    msadmin
    August 11, 2026
  • About
    • Mission
    • Services
    • Contact
  • Shop
    • All Items
    • By Category
    • Cart
  • Newsletter
Font ResizerAa
MedsparkMedspark
Font ResizerAa
  • Home
  • News & Alerts
  • Spotlight
  • Articles
  • About
  • Shop
  • Newsletter
  • Home
  • News & Alerts
  • Spotlight
  • Articles
  • About
    • Mission
    • Services
    • Contact
  • Shop
    • All Items
    • By Category
    • Cart
  • Newsletter
Follow US
Articles

Johns Hopkins finds frontier AI agents fail most complex health tasks

A benchmark test found leading AI agents completed only 28% of healthcare administrative tasks on a first attempt.

MedSpark Staff
By
msadmin
MedSpark Staff
Bymsadmin
Medical, Healthcare, & Biotech/Pharma AI News
Follow:
Published: July 27, 2026
Share
2 Min Read
SHARE

Johns Hopkins Medicine is taking a measured approach to agentic AI, starting with rigorous benchmarking before any large-scale deployment. Early tests found that leading frontier AI agents completed only about 28% of complex healthcare administrative tasks on a first attempt.

Dr. T.Y. Alvin Liu, inaugural director of the Gills AI Innovation Center, said the organization used actAVA’s X-BENCH benchmark to evaluate agents in a simulated environment with 25 healthcare applications and 77 tools. Tasks drew on thousands of pages of managed care policies. Fewer than 8% of agents remained consistently successful across repeated testing.

“The model is necessary but not sufficient,” Liu said. “The harness built for the workflow is what turns a capable model into a deployable agent.” Most failures involved reasoning through policy-rich situations rather than hallucinations or software integration problems.

Johns Hopkins measures reliability first — first-pass completion rates, consistency during handoffs, reasoning failures, and unsafe completions — before calculating financial returns. The health system also built governance structures aligned with the NIST AI Risk Management Framework, HIPAA, and CMS health equity criteria.

For organizations considering agentic AI, Liu recommends defining success across clinical, operational, and financial dimensions, and engaging operational leaders before pilots begin rather than leading with technology alone.

TAGGED:Agentic AIAI AgentsAI benchmarkingAI Governancehealth ITHealthcare AIJohns Hopkins
SOURCES:Healthcare IT News
Share This Article
Facebook Copy Link Print
MedSpark Staff
Bymsadmin
Follow:
Medical, Healthcare, & Biotech/Pharma AI News

You Might Also Like

News & Alerts

Forus triples value to $3B on a $150M Bain-led round

By
msadmin
September 9, 2026
Articles

Embedded transparency is key to equitable AI clinical trials

By
msadmin
July 13, 2026
Articles

Proteomics model picks breast cancer drugs from biopsies

By
msadmin
September 10, 2026
Articles

AI Emerges as a Game-Changer in Humanitarian Healthcare Crisis Response

By
Yu Chi Huang
July 4, 2025

AI news, analysis, and insights for healthcare, biotech, and pharma.

Facebook Twitter Youtube Linkedin
Quick Links
  • News & Alerts
  • Articles
  • Spotlight
  • Events
About Medspark
  • Mission
  • Services
  • Contact

© Copyright 2026 MedSpark. All rights reserved.

Privacy Policy | Legal