FRI finds AI experts underestimated AI progress by years

The Forecasting Research Institute (FRI) has been collecting forecasts on AI progress since mid-2022 through several studies, including its Longitudinal Expert AI Panel (LEAP). The first LEAP round drew 339 experts: 76 computer scientists, 76 industry experts, 68 economists and 119 AI policy specialists. The computer scientists included 30 professors at top-20 institutions and 10 of the 200 most-cited AI authors. The panels also included superforecasters, generalists with a track record of accurate predictions, for comparison.
The widest gap was in math. AI reached gold-medal level at the International Mathematical Olympiad in July 2025, five years before the median expert forecast and ten years before the median superforecaster forecast. Those particular predictions were gathered in 2022, before ChatGPT launched, but FRI says the pattern of underestimation held in later surveys too. AI may also have solved a Millennium Prize Problem, though it remains unclear whether the solution meets the evaluation criteria; in an August-September 2025 survey, experts had put the median odds of such a solution by the end of 2027 at just 10 percent, and superforecasters at 5.4 percent.
In a study of AI capabilities in virology, experts predicted AI models would not match a top team of virologists on a troubleshooting benchmark until 2030; superforecasters said 2034. FRI says that likely happened as early as April 2025. A cybersecurity benchmark showed similarly conservative forecasts. Economic predictions were also too low: experts put the median for the highest annual recurring revenue (ARR) of any AI company at the end of 2026 at $20 billion, economists said $16 billion and superforecasters said $25 billion. FRI cites a figure of roughly $100 billion for Anthropic's ARR in September 2026 as likely already reached.
Not every forecast ran too low, though. Biosecurity experts predicted that 22.5 percent of participants using a language model would complete biological lab tasks; virologists expected 40 percent and superforecasters 16.2 percent. In a controlled trial, only 5.2 percent succeeded with a language model and internet access, compared with 6.6 percent using the internet alone, and FRI describes the language model as making no measurable difference, though it notes the trial was small. Experts may also have overshot on self-driving cars: their median forecast for the share of autonomous US ride-hailing trips in 2027 was 7.3 percent, versus an LLM projection of 2.5 percent. FRI says forecasts on economic growth, employment and major AI harms cannot yet be reliably judged.
At the same time, respondents are revising expectations upward: among those who completed two surveys nine months apart, the average probability assigned to AI becoming a "technology of the century" rose from 31 to 36 percent for experts and from 28 to 35 percent for superforecasters.
Going forward, FRI plans to highlight a subsample of respondents who expect very rapid AI progress through 2040 and to publish continuously updated LLM forecasts alongside the human ones; according to ForecastBench, some models already match superforecasters on certain question types. FRI also wants to identify the most accurate LEAP panelists and feature their forecasts once enough data is in. The institute flags a limitation in its own data: underestimates become visible as soon as reality overtakes a prediction, while overestimates only become clear once a deadline passes, which tilts the interim report toward finding cases of excessive caution. FRI also notes that some of its own assessments rely on LLM projections that use information the original human forecasters did not have.
Key facts
- AI reached gold-medal level at the International Mathematical Olympiad in July 2025, five years before the median expert forecast and ten years before the median superforecaster forecast
- Experts gave only 10 percent median odds (superforecasters 5.4 percent) that AI would solve a Millennium Prize Problem by end of 2027, in an August-September 2025 survey, yet FRI says AI may already have done so
- Experts expected AI to match a top team of virologists on a troubleshooting benchmark only by 2030 (superforecasters: 2034); FRI says that likely happened by April 2025
- Economic forecasts were also too conservative: the median expert forecast for the highest AI-company ARR by end of 2026 was $20 billion, versus a roughly $100 billion figure FRI cites for Anthropic as of September 2026
- Not all forecasts ran low: a controlled trial found only 5.2 percent succeeded at biological lab tasks with a language model and internet access, versus 6.6 percent with internet alone, well under experts' predicted 22.5 percent
Why it matters
FRI's Longitudinal Expert AI Panel is a systematic attempt to track how well domain experts and top forecasters predict AI's trajectory rather than relying on anecdotal impressions. Its interim finding, that panels of computer scientists, economists, policy specialists and superforecasters missed several capability milestones by five to ten years, matters for anyone who leans on expert timelines to plan policy, investment or safety work around AI.
Who it affects
The finding bears on AI policy specialists, economists modeling AI-driven growth, biosecurity and cybersecurity risk assessors, and the wider forecasting and AI-safety communities that rely on expert elicitation to set timelines for capability milestones.
How to use it
This is a research report rather than a product, so the practical takeaway is a calibration lesson: discount expert and even superforecaster timelines that assume slow AI progress on capability benchmarks, but do not swing to the opposite extreme, since the same report shows adoption and real-world-impact forecasts (biosecurity task completion, self-driving ride-hailing share) were sometimes too high rather than too low.
How solid is it
FRI has run this tracking since mid-2022, and the first LEAP round alone drew 339 experts across four groups plus a superforecaster panel, with sourcing details (institution tier, citation rank) given for the computer scientists. The source describes the report only as an 'interim report' without naming a publication date, authors or the full title. FRI itself flags a structural bias in the data: underestimates become obvious as soon as reality overtakes a prediction, while overestimates only surface once a deadline passes, which tilts an interim report toward finding cases of excessive caution; FRI also notes some of its own assessments use LLM projections built on information the original forecasters lacked.
Risks and caveats
The report does not identify which specific AI model or company reached IMO gold-medal level, solved the Millennium Prize Problem, or matched virologists and cybersecurity experts, apart from citing Anthropic for the ARR figure. FRI itself says forecasts on economic growth, employment and major AI harms cannot yet be reliably judged, and the biosecurity trial that found no measurable uplift from language-model use was small.