OpenAI Foundation funds grants to tackle AI's biology data gap

OpenAI Foundation funds grants to tackle AI's biology data gap

The OpenAI Foundation, OpenAI's nonprofit parent, has announced its first round of grants for an initiative meant to supply artificial intelligence with more biological and medical data. The article's own subheading calls the effort 'Data for Public Health,' while the body twice describes it as 'Public Data for Health'; the two names conflict and the piece never explains why. The foundation says it will pay to create 'high-quality scientific datasets' because, in the words of Morgan Levine, a former vice president for computation at the longevity company Altos Labs, 'everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology.' The foundation's own statement frames the goal as pairing more capable models with more real-world observations to produce breakthroughs in preventing and curing disease.

The idea behind the largest strand of the initiative came from Ruxandra Teslo, a policy analyst who focuses on clinical trials and who also writes for Works in Progress and serves as a nonresident fellow at the Institute for Progress, a Washington think tank. Last year she proposed bidding at the bankruptcy proceedings of failed biotech companies to obtain detailed regulatory filings, manufacturing strategies and safety data, information usually guarded as trade secrets. She called this material 'biotech's lost archive' and argued it could help train AI systems to act as copilots in the drug approval process. The OpenAI Foundation has now granted $500,000 to pursue the idea through 1Day Sooner, an advocacy group for clinical trial volunteers that Teslo advises.

1Day Sooner's president and cofounder, Josh Morrison, says the grant will help the group prove it can obtain the data troves of bankrupt companies; he estimates that nonexclusive copies of a company's dataset could cost only 'a few tens of thousands of dollars' each. The group already holds three datasets, two of them donated by Lumen Bioscience, a biotech that had itself previously used the Chapter 11 process to gain insight into a competitor's drug development. Two other attempts this year to obtain drug company files fell through when 1Day Sooner's bids were not accepted. The files it is after are known as common technical documents: the back-and-forth between companies and regulators plus detailed scientific and medical measurements that together cover almost everything known about a drug. Teslo argues a stockpile of these documents could turn an AI into a regulatory expert, which she sees as one of the more realistic ways AI could help speed cures to market, since, in her view, about 70% of the money and time in drug development goes into clinical development, a process she calls a black box, especially for small biotech companies.

Bankruptcy filings are becoming what some call a 'new land grab' for AI training data. Last month Google won a bid for the corporate data of the failed carrier Spirit Airlines, including 100 million emails, a deal that drew objections from flight attendants and others worried that private or proprietary information could be exposed.

Alongside the 1Day Sooner grant, the foundation's initial round of data grants includes $40 million toward a University of North Carolina, Chapel Hill program that collects data on novel cancer vaccines, plus unspecified support for OpenAdmet, a group that runs competitions in which researchers try to predict drug effects; the article does not break down how the $40 million splits between the two. Separately, the foundation's largest single gift to date, $100 million, went in August to the Common Health Coalition, which helps patients get access to hepatitis C drugs. Jacob Trefethen, an executive at the foundation, says it operates essentially separately from OpenAI but shares its mission of ensuring AI 'benefits all of humanity,' and that the foundation hopes to give away $1 billion by the end of the year.

The foundation can draw on unusual resources: it holds a 26% equity stake in OpenAI, which the article says could leave it sitting on some $250 billion in stock value, putting it on track to become the richest charitable organization on the planet (for comparison, the Gates Foundation and an associated trust held about $180 billion at the end of 2025). Even so, the San Francisco-based foundation is still hiring for many key roles and only started ramping up its grantmaking this year. The announcement arrives alongside two larger backdrops the article does not tie together: OpenAI's Sam Altman has restructured the company from a nonprofit into a for-profit corporation that is now planning an initial public offering which could value it at $1 trillion, and unnamed AI company insiders have put the chance of human extinction within the next decade at 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to slow the pace at which AI capabilities improve, so that risk prevention work can catch up.

Key facts

  • The OpenAI Foundation's first round of data grants includes $500,000 to 1Day Sooner, to bid at bankrupt biotech companies' proceedings for regulatory filings, manufacturing strategies and safety data normally kept as trade secrets.
  • It also includes $40 million toward a University of North Carolina, Chapel Hill program collecting novel cancer vaccine data, plus unspecified support for OpenAdmet; the article does not say how the $40 million splits between them.
  • 1Day Sooner already holds three datasets, two donated by Lumen Bioscience, and its president Josh Morrison estimates nonexclusive copies of a bankrupt biotech's data could cost only 'a few tens of thousands of dollars' each; two other bid attempts this year failed.
  • The foundation holds a 26% equity stake in OpenAI, potentially worth $250 billion, versus the about $180 billion the Gates Foundation and its trust held at the end of 2025; it aims to give away $1 billion by year end, per executive Jacob Trefethen.
  • Google set a precedent last month by winning bankrupt carrier Spirit Airlines' corporate data, including 100 million emails, drawing objections from flight attendants over exposed private data.

Why it matters

AI systems built for medicine are limited less by model design than by what they have to learn from. The foundation's own framing, echoed by Morgan Levine, is that data, not model capability, is now the biggest bottleneck to applying AI in biology. Regulatory filings, manufacturing strategies and safety data from failed drug companies have historically stayed locked away as trade secrets even after the companies collapse; funding an effort to legally acquire that material at bankruptcy sales opens a category of information that has never fed an AI model before. The bet, in Teslo's telling, is that documents describing how real drugs succeeded or failed in the approval process, not published research alone, could make an AI useful as a genuine copilot for navigating drug regulation.

Who it affects

Direct beneficiaries include 1Day Sooner, the University of North Carolina, Chapel Hill cancer vaccine program, and OpenAdmet. Failed biotech companies become unwitting data sources through their bankruptcy filings, and Lumen Bioscience, which has already used this same route itself, shows the tactic is not unique to AI funders. Drug developers, regulators and eventually patients are the intended long-term beneficiaries if the data does help produce a genuine 'regulatory expert' AI, as Teslo describes it. The Spirit Airlines precedent, where Google acquired 100 million emails from the failed airline's bankruptcy and drew objections from flight attendants, shows the same mechanism can expose people who never agreed to have their data used this way, a risk that extends to employees or trial participants named in a biotech company's bankruptcy files.

How to use it

There is no product to install here, but two access points already exist. OpenAdmet runs open competitions that researchers can enter now to work on predicting drug effects. The foundation says it hopes to give away $1 billion in grants by the end of the year and describes its grantmaking, per Trefethen, as directed at 'external nonprofits, research institutions, and other third parties,' so further calls for proposals are the likely next step for anyone working on biomedical datasets. Morrison's estimate that a nonexclusive copy of a bankrupt biotech's dataset costs only a few tens of thousands of dollars suggests the resulting archives, if the bids succeed, could eventually reach researchers beyond the foundation's own grantees, though the article does not say whether or how that would happen.

How solid is it

The reporting rests on named, on-the-record sources: Morgan Levine, Jacob Trefethen, Josh Morrison and Ruxandra Teslo are all quoted directly, and the dollar figures for the individual grants, $500,000 and $40 million, are stated plainly rather than estimated. Some of the larger, more dramatic figures are explicitly hedged rather than confirmed: the $1 trillion IPO valuation is 'could value,' the $250 billion in potential foundation stock value is 'potentially sitting on,' and Morrison's tens-of-thousands-of-dollars figure is his own estimate, not a completed transaction. The piece also has an internal inconsistency worth flagging: its own subheading names the initiative 'Data for Public Health,' while the body twice calls it 'Public Data for Health,' and it never says which is correct. None of the events carry an absolute calendar date either; 'last year,' 'last week,' 'last month,' 'this year' and 'by the end of the year' are all relative to the article's own publication date, so it is not even clear which year the August gift or the year-end goal refers to. The article also does not specify whether the $40 million grant covers the University of North Carolina program alone or is split with OpenAdmet, does not name which AI insiders back the 10% extinction-risk estimate, and does not explain when or how the foundation acquired its 26% stake in OpenAI.

Risks and caveats

Mining bankruptcy filings for data raises the same consent questions the Spirit Airlines case already surfaced: the employees, trial participants and business partners named in a failed biotech's regulatory and safety filings did not agree to have that information reused to train AI, only to have their employer or drug developer hold it. The 'trade secret' framing cuts both ways: the material has stayed private for reasons beyond commercial value, including safety data tied to real trial participants, and the article gives no indication that anyone screens the acquired files for personal information before they reach a model. There is also a wider tension the article does not resolve: last week OpenAI's Sam Altman and xAI's Elon Musk endorsed Anthropic CEO Dario Amodei's call to slow AI capability development over extinction fears some insiders put at 10% or more, while today's announcement instead funds new data meant to make AI more capable in a high-stakes domain. Finally, the article names only four grant or gift recipients in total, 1Day Sooner, the University of North Carolina program, OpenAdmet and the Common Health Coalition, and does not say how many grants the initial data round included beyond these, so it is unclear how much of the foundation's stated $1 billion year-end goal these figures represent.

“Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology.”

— Morgan Levine, former vice president for computation at Altos Labs