Stanford Medicine leaders say benchmark scores are not enough to monitor large language models once clinicians actually start using them, based on lessons from deploying the hospital’s ChatEHR system.
In a Comment published in Nature Medicine on August 10, Stanford’s Nigam Shah and co-authors Michael Pfeffer, Niraj Sehgal, and Euan Ashley describe what happened when the LLM-driven tool, which lets clinicians ask questions about patient records, moved from pilot to production at one of the country’s largest medical centers.
Their core finding: benchmark-based evaluations are insufficient for monitoring and evaluating interactions driven by clinicians, and hospitals need new methods for tracking performance in real use. Off-the-shelf benchmarks measure how a model answers fixed questions, but they miss how clinicians actually prompt the system, how responses change over time, and when answers drift into unsafe territory.
ChatEHR, developed with Stanford’s clinical AI team, gives doctors a natural-language window into medical records. The deployment experience suggests the hardest part of clinical AI is not building the model but watching it after launch.
The authors call for monitoring approaches designed around clinician behavior, including ongoing evaluation of real interactions rather than relying on pre-deployment test scores. For health systems racing to deploy LLMs, the Comment is a reminder that a model’s performance on a test set says little about its performance at the bedside.
