A $1.29M research contract from the FDA will pay for a new method of judging radiology reports that AI writes. The recipient is Cognita Imaging, a subsidiary of Mosaic Clinical Technologies.
The problem is scale. Reader studies remain the gold standard for checking a medical AI system, but expert review cannot cover the range of cases a generative model will meet in practice. Cognita’s answer is what it calls LLMs-as-a-jury: several language models independently grade human-drafted and AI-generated reports, and radiologists adjudicate the disagreements that matter clinically.
Building and validating the framework is step one. Step two pushes it onto a scale reader studies cannot reach: roughly 1 million patient exams from a large, diverse US cohort. Patient groups, care settings, imaging equipment and rare findings all get checked. A second pass rebuilds smaller cohorts to show what a trimmed study would miss.
The project is led by Cognita co-founder Akshay Chaudhari, who teaches radiology at Stanford. Once a model starts writing the whole report, he said, the evaluation problem itself changes. A few hundred cases can show whether a system works narrowly, but not every way it can fail in practice.
Deliverables include software code, comparisons of large and small validation cohorts, and discrepancy studies reviewed by radiologists. Cognita will also hand over guidance for building jury frameworks. The work builds on GREEN, its open-source tool for flagging clinically meaningful differences between two reports.
Why it matters: generative reporting tools are arriving faster than the methods used to check them, and automated review may be the only way to keep pace.
