Stanford Medicine reports that benchmark tests cannot monitor real clinical LLM use, based on its ChatEHR rollout.