A preprint estimates that roughly nine in ten biomedical papers from December 2025 show signs of AI-assisted writing.
A scoping review maps five categories of mental health harms tied to LLM-based chatbots.
Stanford Medicine reports that benchmark tests cannot monitor real clinical LLM use, based on its ChatEHR rollout.
A clinically validated framework ran 810 conversations across nine chatbots and found supportive replies can reinforce vulnerable users' harmful thinking.
GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 outperformed dedicated clinical AI systems in head-to-head testing across 1,100 medical questions.