Peer reviewers flagged fabricated authors, the papers got accepted as orals anyway

Two peer reviewers who write under the names Caleb and Isaac reviewed 22 paper submissions this summer across NeurIPS, WACV and TerraBytes, a geospatial workshop at ECCV. Fifteen of the 22, 68%, contained entirely fabricated citations, fabricated author lists for existing papers, or writing that was clearly LLM-generated. Caleb reviewed 11 papers and found 5 with hallucinated authors or entirely fabricated bibliography entries. Isaac also reviewed 11: two NeurIPS position papers, five for TerraBytes and four for WACV, and found problems in 9 of the 11. Four of the five TerraBytes assignments had fabricated citations or authors.

Isaac's worst find: two submissions cited papers whose real authors Isaac knew personally, with the actual names swapped out for invented researchers close enough to survive a skim. Isaac recommended rejection and flagged both papers to the organizers directly. Both were accepted for oral presentation anyway, on condition the authors simply fix the hallucinated references. Caleb's worst cases were a 53-page NeurIPS paper so dense with hallucinated jargon that checking its bibliography felt pointless, a 40-page paper with similarly invented references, and a WACV submission that otherwise read well but listed the authors of the SatMAE paper as Yuyang Cong, Saurabh Khanna and Chen Meng, when the real SatMAE author list starts with Yezhen Cong, Samar Khanna and Chenlin Meng (arXiv:2207.08051).

The two reviewers say this fits a pattern documented elsewhere. A Nature analysis from April found at least tens of thousands of 2025 publications probably contain invalid AI-generated references. An audit of arXiv, bioRxiv, SSRN and PubMed Central by Zhao et al. estimated roughly 146,900 hallucinated citations across 2025 publications, spread thinly rather than concentrated among a few bad actors, with early-career researchers and small teams the most likely to include them; tracing bioRxiv preprints to their published versions, the audit found 85.3% of the hallucinations remained uncorrected. An audit in The Lancet covering 2.5 million biomedical papers found the share of papers with at least one fabricated reference rose six-fold in two years: from one in 2,828 papers in 2023 to one in 458 in 2025 and one in 277 in early 2026. Ansari (2026) traced 100 hallucinated citations that appeared in papers actually published at NeurIPS 2025; every one passed three to five expert reviewers, and the 53 papers carrying them, about 1% of that year's acceptances, remain in the proceedings.

The problem runs both ways: reviewers are using AI too. Pangram, a company that sells an AI writing detector, analyzed ICLR 2026 reviews and found 21% of them, 15,899 reviews, were fully AI-generated, with over half showing some AI involvement. Gartenberg et al. (2026) measured a 42% post-ChatGPT surge in submissions to the journal Organization Science and found over 30% of its peer reviews now use some degree of AI. ICML 2026 planted prompt-injection stings inside submissions and caught 795 reviews, about 1% of all reviews, written by 506 reviewers who had been assigned the no-LLM policy. Li et al. (2026) showed the reverse exploit works too: adversarially rewritten abstracts improved AI-generated review outcomes without changing the paper's actual scientific content, succeeding about 38% of the time in their strongest attack and inflating acceptance ratings on a 10-point scale by 1.31 points for Gemini 3 Flash reviewers and 0.88 points for GPT 5.4 Mini reviewers.

Caleb and Isaac say their reviewing habits have changed: Caleb now checks the bibliography right after reading the abstract, and Isaac built a bib-audit skill they both use, though Isaac still verifies each reference by hand against Google Scholar or the relevant database. Their top tells for LLM-written text: stock phrases, heavy bold formatting and dashes; dense, hard-to-parse sentences; overhyped results; and, per Isaac, numbers in the prose that do not match the paper's own tables. Isaac argues papers with hallucinated references should simply be desk rejected, since the authors clearly did not spend enough time with their own work to expect a reviewer to. Both say they use Claude daily for drafting and literature review on their own writing, including this post, but insist they verify, edit and rewrite everything themselves rather than submitting raw model output.

Key facts

  • 15 of the 22 papers, 68%, that Caleb and Isaac reviewed this summer for NeurIPS, WACV and TerraBytes had fabricated citations, invented authors or clearly LLM-generated writing.
  • Two submissions with invented co-authors, close enough to real names to survive a skim, were accepted as oral presentations after Isaac flagged them and recommended rejection.
  • A Lancet audit of 2.5 million biomedical papers found the fabricated-reference rate rose six-fold in two years, from one in 2,828 papers in 2023 to one in 458 in 2025, reaching one in 277 by early 2026.
  • Pangram found 21% of ICLR 2026 reviews, 15,899 of them, were fully AI-generated, and ICML 2026's prompt-injection sting caught 795 reviews from 506 reviewers barred from using LLMs.
  • Li et al. (2026) showed adversarially rewritten abstracts can game AI reviewers, inflating acceptance scores by up to 1.31 points on a 10-point scale without changing the paper's actual content.

Why it matters

Peer review is the mechanism that is supposed to catch bad science before it enters the record, and this account shows it failing at the most basic level: reviewers who personally knew the real authors of a cited paper still could not stop invented co-authors from making it into an oral presentation. The scale is not anecdotal. The independent audits cited in the post, from Nature, Zhao et al. and The Lancet, all point the same direction: fabricated references in published papers have grown sharply since 2023, and the review process is not catching most of them before or after publication.

Who it affects

Volunteer peer reviewers carry the direct cost, spending unpaid hours on submissions that turn out to be unreadable or fraudulent, with some conferences now making review mandatory for anyone who submits. Early-career researchers and small teams are, per Zhao et al., the group most likely to include hallucinated citations. Conference organizers at NeurIPS, WACV, ICLR and ICML, and journals like Organization Science, face a credibility problem as both submissions and reviews increasingly involve undisclosed AI use. Readers of the eventual proceedings inherit papers that already carry fabricated references, such as the 53 NeurIPS 2025 papers Ansari traced.

How to use it

The post's practical advice, from two reviewers who do this work regularly: check the bibliography immediately after the abstract, before investing time in the rest of the paper; verify suspicious author lists manually against Google Scholar or the venue's own database rather than trusting a citation tool alone; and watch for tells beyond citations, prose that restates table numbers without adding anything, overhyped gains under 1% reported with no mean or standard deviation, and fancy math where a citation would do. Isaac's bib-audit skill, referenced but not detailed in the post, automates part of the citation check.

How solid is it

This is a firsthand account from two anonymous peer reviewers on their own blog, not itself peer-reviewed, so the central anecdote, the two papers with swapped-in fake authors accepted as orals, cannot be independently checked: neither the venue nor the paper titles are named. The surrounding context draws on several external studies, a Nature analysis, an audit by Zhao et al., a Lancet audit of 2.5 million papers, work by Ansari, Pangram's ICLR analysis, Gartenberg et al. on Organization Science, ICML's own prompt-injection sting, and Li et al. on adversarial abstracts, several of them peer-reviewed or from established outlets, and their numbers point in the same direction independently of each other.

Risks and caveats

Neither Caleb's nor Isaac's full name, institution or affiliation is disclosed, and the venue of the two accepted-with-fake-authors papers is not named, so the specific claim rests on trust in the authors' own account. The post does not say whether the hallucinated references were actually corrected after acceptance, only that correction was the stated condition. The cited studies use different definitions and methods for what counts as a hallucinated citation or an AI-generated review, so their percentages are not directly comparable to each other or to Caleb and Isaac's own 68% figure.

“Papers with hallucinated references should just be desk rejected as they are clearly not ready for acceptance.”

— Isaac, co-author of the post