Stanford and Harvard researchers debut MAST v1.0, a new benchmark framework for measuring clinical AI performance across real healthcare tasks.