A new benchmark from Synthio Labs finds that voice AI can sound convincingly human and still botch the name of a newly approved medicine.
The benchmark is called DOSE, for Drug-name Oral Synthesis Evaluation. It put nine commercial text-to-speech systems through 274 medicine names, 146 of them recent approvals, each spoken inside a clinical sentence. Grading ran from 0 to 5 against verified reference pronunciations, and 4 or above was a pass.
The drop was steep and universal. ElevenLabs’ eleven_v3 gave up 26 points of accuracy once the names turned new, and Microsoft Azure failed most generic names outright. Another system read Xofluza aloud as separate characters.
First place went to Synthio’s own RxPronounce model. It passed 91.2% of all names and 87.0% of the new ones, 10.9 points clear of the runner-up. The company sells voice and agentic AI to pharmaceutical clients and put the dataset and audio on Hugging Face.
According to Synthio co-founder and CEO Supreet Deshpande, the failures track where training data thins out. A legacy drug like metformin gives models no trouble. The molecule cleared three months ago is a different story, and it is precisely the name a launch team or a patient needs said correctly.
Why it matters: the WHO, the ISMP and the FDA all treat sound-alike drug names as a safety category of their own. Voice agents introduce one more speaker into that problem, and until now nobody had quantified how often they get the names wrong.
