Voice researchers should stop building one-off classifiers and instead develop a shared voice-biomarker foundation model for tracking ALS and screening Parkinson’s disease, argue the authors of a new npj Digital Medicine perspective published September 4.
Speech production leans on respiratory, laryngeal, articulatory and executive control, so neurodegeneration can leave acoustic traces before standard clinical scales move. Yet the field stays fragmented into small models built for a single condition and language, many of which fail when tested on speakers they never met during training. The authors also warn about data leakage, where recordings from the same person leak across training and test sets and flatter results.
Their proposal is one self-supervised backbone pretrained on large, diverse, ethically sourced speech and then adapted to specific jobs. ALS bulbar-progression monitoring looks like the strongest first use case because speech measures can respond faster than the coarse ALSFRS-R scale. One ALS speech analytics platform already carries FDA breakthrough device designation, though that status is not marketing approval.
Parkinson’s screening serves as the cautionary tale, with modest real-world accuracy and a case for favoring articulation-rich tasks over simple phonation checks. No voice-derived endpoint for either disease has been qualified by the FDA or EMA, so the paper sets minimum methodological and governance standards and a roadmap for prospective validation before any model earns clinical trust.
