Stanford and Harvard researchers have launched the Medical AI Superintelligence Test, or MAST, a new framework designed to track how clinical AI models perform across meaningful healthcare tasks.
Published in Nature Medicine by the ARISE Healthcare Network, MAST v1.0 measures AI models across clinical domains like diagnosis, management, and safety, as well as general domains including radiology and multimodal reasoning. The goal is a shared infrastructure for rapid, high-quality AI benchmarking.
The researchers found that no single frontier model dominates. “The highest-performing models are closely clustered on the composite rankings, but the ordering changes across clinical dimensions,” the team reported. “There is no single model that dominates.”
ARISE, which stands for AI Research and Science Evaluation, was established by Stanford in 2024 as a collaborative network of academic medical centers. The group argues that existing benchmarks are misleading and insufficient for assessing AI safety in real clinical settings.
“A model can perform well on knowledge questions while failing to integrate that knowledge in complex clinical scenarios,” ARISE explained. “It can make the right diagnosis while recommending an unsafe plan.”
The launch version maintains evaluations with standardized model runs and public reporting across diagnosis, management, and other domains. Future iterations will cross-benchmark trait analysis and use evaluations as probes of underlying clinical behaviors, since traits like aggressiveness cannot be captured by a single test.
The framework builds on the NOHARM study, which evaluated clinical AI safety across 45 large language models and four clinical AI systems. ARPA-H recently awarded $3.8 million to Stanford and Beth Israel Deaconess researchers under its PACT project to build task-first AI benchmarks grounded in electronic health record data.
