An independent audit of OpenEvidence’s AI answers found no fabricated citations across nearly 5,000 references. The clean record stands out at a time when invented sources are surging across the biomedical literature.
The study appeared September 2 in npj Health Systems. Investigators ran 150 standardized clinical prompts through the retrieval-augmented platform across oncology, cardiology, rheumatology, psychiatry and infectious disease. They then verified all 4,979 returned citations against PubMed and original databases. None was invented. Three ASCO meeting abstracts carried author-attribution errors, and no predatory journals turned up.
The reference lists skewed recent and respected. Median publication year was 2022. Roughly 70% of citations dated from 2020 or later, nearly half came from Q1 journals, and about a quarter were flagship titles such as NEJM, JAMA and The Lancet.
The outcome matters because hallucinated citations are endemic in general-purpose LLMs. Earlier work estimated that 30-69% of AI-generated biomedical references are fabricated, and a Lancet audit of 2.5 million published papers found fabrication rising 12-fold since 2023, touching one in every 277 papers.
OpenEvidence restricts retrieval to licensed, indexed sources such as JAMA Network, Cochrane and NCCN content, which the authors say explains the result. They caution the audit confirms references exist, not that they support the platform’s clinical claims. Researchers from Tel Aviv University’s Gray Faculty of Medical and Health Sciences and the BRIDGE GenAI Lab at Beth Israel Deaconess Medical Center led the work. It lands as OpenEvidence unveils a new family of AI models for clinicians, first reported by STAT News on September 3.
