AI unicorns barely publish scientific research, study finds
A preprint posted on 16 July on bioRxiv looked at 317 AI unicorns, private companies valued at more than $1 billion, founded between 1998 and 2025, to see how much they contribute to the published scientific record. It found that more than half of these firms have never played a leading role in publishing a peer-reviewed paper or preprint. Collectively, the unicorns accounted for only one in every 1,000 AI papers published in 2025, despite their prominence in the industry.
Where research did appear, it was heavily concentrated in a handful of names. The top 5% of firms account for more than 90% of all citations in the sample the researchers built. OpenAI alone was responsible for nearly 40% of all citations in the dataset, followed by the Chinese computer vision company Megvii and the platform Hugging Face. The full dataset assembled for the study comprised 2,077 publications: 1,389 peer-reviewed papers and 688 preprints.
The imbalance shows up inside individual companies too. OpenAI employs roughly 4,500 people, yet only eight of its researchers had authored five or more of the qualifying papers, a striking gap between headcount and public research output.
Co-author John Ioannidis, a metascientist at Stanford University, called the finding a paradox for a field that presents itself as reshaping science: rather than documenting discoveries so other researchers can evaluate and build on them, most AI unicorns leave almost no public trace of their internal research.
Key facts
- More than half of the 317 AI unicorns studied have never played a leading role in publishing a scientific paper or preprint.
- Collectively, AI unicorns produced just one in every 1,000 AI papers published in 2025.
- The top 5% of firms account for over 90% of citations in the sample; OpenAI alone accounts for nearly 40%, followed by Megvii and Hugging Face.
- OpenAI has about 4,500 employees, but only eight of them authored five or more qualifying papers.
- The findings come from a preprint posted 16 July on bioRxiv, co-authored by Stanford metascientist John Ioannidis.
Why it matters
AI companies routinely present themselves as pushing the frontier of science, yet the study suggests most of the industry's most valuable private firms contribute almost nothing to the public scientific literature that lets outside researchers check, replicate or build on their work. The irony the study points to is that these firms largely built their models on published academic and non-academic data, but do not reciprocate by publishing their own findings for others to use.
Who it affects
It affects the wider research community, which loses visibility into what advanced techniques actually work inside frontier labs; competitors, who can no longer learn from published methods the way earlier generations of AI research operated; and policymakers and journalists who rely on the public record to gauge how much scientific progress is really happening inside heavily funded private companies.
How to use it
Treat claims that a given AI startup is advancing science with more scrutiny: this dataset gives a concrete way to check whether a company has an actual publication record behind its marketing, rather than taking frontier-lab branding at face value. Researchers weighing collaboration or citation of a company's work can use the concentration figures (OpenAI, Megvii and Hugging Face account for most of the citations) to gauge where genuine open research activity is actually happening.
How solid is it
The finding is quantitative and specific: a defined set of 317 unicorns, a dataset of 2,077 publications, and clear concentration statistics (over 90% of citations in the top 5% of firms). It carries the standard caveat of any bioRxiv preprint: it has been posted for public scrutiny but has not yet completed formal peer review.
Risks and caveats
As a preprint, the methodology, including how 'unicorn,' 'leading role' in a paper, and 'qualifying paper' were defined, has not been through independent peer review and could shift on revision. The count also only captures formal papers and preprints; it does not measure other channels companies use to share technical work, such as blog posts, technical reports or open-source model releases, which some firms may treat as their primary form of disclosure instead of the academic literature.
“For a field that is supposedly reshaping science and is so advanced in terms of scientific potential, not having any scientific documentation seems like a very weird paradox”
— John Ioannidis, metascientist at Stanford University, study co-author