OpenAI’s nonprofit parent is paying to build the scientific datasets it says medical AI is missing. The program is called Data for Public Health.
Its logic is simple. Models cannot uncover cures if they never see enough observations. Morgan Levine, who previously led computation at the longevity company Altos Labs, said the field now broadly agrees on where the constraint sits: data, not architecture, is what limits biology.
A first round of grants is already committed. The University of North Carolina at Chapel Hill receives $40M to gather information on novel cancer vaccines. OpenAdmet, which runs contests where researchers forecast drug behavior, also gets support.
One smaller award stands out. Policy analyst Ruxandra Teslo proposed bidding at failed biotech companies’ bankruptcy proceedings to obtain regulatory filings, manufacturing plans and safety results she calls biotech’s lost archive. The foundation put $500K behind the idea, and the advocacy group 1Day Sooner will execute it.
Money is not the constraint here. The foundation owns a 26% stake in OpenAI, which is preparing an IPO. That position could be worth roughly $250B, which would make it the wealthiest charity in the world. Its largest single grant to date was $100M in August to the Common Health Coalition for hepatitis C drug access.
The program arrives as biology data turns into a competitive front. Labs are racing to train foundation models on cellular and protein information, and whoever assembles the best datasets may set the pace.
