Stanford and Harvard researchers debut MAST v1.0, a new benchmark framework for measuring clinical AI performance across real healthcare tasks.
The White House, FDA, and ONC launch a month-long effort to create consensus principles for clinical AI benchmarking.
A benchmark test found leading AI agents completed only 28% of healthcare administrative tasks on a first attempt.