Perplexity cites 215,128 AI-only pages from apparent content farms

A report published at trellner.com tested how Perplexity's web-grounded models pick software recommendations. On 2 September 2026 the researchers sent 380 buyer-intent categories, from "CRM software" to "museum collection management software," to perplexity/sonar and perplexity/sonar-pro through OpenRouter, one prompt per category per model, 760 calls total, asking each for a ranked top five with each product's official homepage domain. Categories were written before any results were seen. Google was excluded because grounding a Gemini model on OpenRouter routes it through OpenRouter's own web-search plugin rather than Google's own retrieval, so the findings describe Perplexity only.

The run produced 3,800 recommendation slots naming 1,807 distinct products and 7,534 citations across 2,055 distinct domains. Checked against the Tranco top-1M list for 2026-09-01, 59.8% of citations pointed to domains ranked worse than #100,000, and 23.4% pointed to domains absent from the top million entirely. Among the 5,768 citations that did resolve to a ranked domain, the median Tranco rank was 71,611. The ten most-cited domains took only 17.3% of all citations, so the pattern is not a few famous sites dominating; 751 of the 2,055 cited domains, 36.5%, do not appear in the Tranco top million at all, and those unranked domains are newer on average (median first Wayback capture in 2020, versus 2011 for ranked domains; 16.6% of unranked archived domains were first captured in 2025 or later, against 1.6% of ranked ones). Wikipedia was cited three times out of 7,534.

The third-most-cited domain overall, ahead of Gartner, is guideflow.com, a vendor of interactive product demos that does not compete in any of the 380 categories tested. Its marketing blog was cited 194 times across 96 of the 380 categories, using 96 distinct blog URLs (one per category, six of them an Estonian-locale duplicate) drawn from a sitemap listing 3,351 blog URLs and 2,176 distinct posts. The report stresses this is not deceptive: Guideflow runs an ordinary large content-marketing blog, and the finding is about what Perplexity's retrieval does with it, not about Guideflow's conduct.

A separate cluster looks more deliberately built for AI retrieval. Three sites, wifitalents.com (71 citations, 27 categories), worldmetrics.org (60 citations, 22 categories) and gitnux.org (50 citations, 23 categories), together account for 181 citations (2.4% of the total) across 41 categories, and appear to be one operation: all three were registered through NameCheap between December 2023 and May 2024, none existed before December 2023, all three delegate DNS to the same two Cloudflare nameservers, and all three run an identical page template and navigation. Each keeps exactly six blog posts, all eighteen of which write about the other brands in the set, plus a fourth site on the same nameserver pair, zipdo.co, whose homepage carries the same title format. Their sitemaps list 103,578, 107,083 and 105,541 URLs respectively, of which 70,731, 71,684 and 72,713 are generated /best/-software/ pages: 215,128 such pages across the three sites combined, against six real blog posts each, even though there are nowhere near 215,128 actual software categories. Fetched on 2 September 2026, worldmetrics.org and gitnux.org each returned an HTML title of "Facts & Grounding Page" prefixed with their own brand name, and near-identical meta descriptions describing an "independent market research company" offering a "machine-readable record" of "verified facts." worldmetrics.org also advertises paid custom research starting at €5,000, ready-made reports from €499, and vendor selection from €2,500. Fetching the same "project estimation software" page from all three brands, the report found Gitnux's top pick absent from Worldmetrics' top five, each page crediting three different named staff (nine distinct people total across the three sites for one question), and an unrendered template placeholder reading "Within the next 26 days" on two of the pages and "Within the next 40 days" on the third.

Checking the 1,502 distinct vendor homepages the models supplied, the report found 17 domains (1.1%) dead or unreachable, including graphiql.com, todo.com and aquasecurity.io, and another 92 (6.1%) redirecting to a different registrable domain, mostly ordinary rebrands. Two redirects stood out: asked for research data management platforms, sonar-pro correctly named datadryad.org while sonar named dryad.co, which redirects to an Indonesian gambling site; asked for data quality tools, sonar named montecarlodata.com while sonar-pro named montecarlo.com, which redirects to the Monte Carlo casino group in Monaco.

The report is explicit about its limits: sonar and sonar-pro returned byte-identical citation lists in 289 of 380 categories, with a 0.898 Jaccard overlap and the same top pick in 290 of 380 categories, meaning the two tiers largely share one retrieval layer rather than acting as independent checks; only Perplexity was measured, with no claim made about ChatGPT, Gemini, Copilot or Google's AI Mode; common control of the three sites is inferred from shared infrastructure and template, not proven, and none of the sites names an owner; and the report did not test whether removing the flagged sources would actually change Perplexity's recommendations. The full dataset, citations, Tranco and Wayback lookups, vendor liveness checks and the scripts used are published under CC BY 4.0.

Key facts

  • Across 380 software categories and 7,534 citations from Perplexity's sonar and sonar-pro models, 59.8% pointed to domains ranked worse than #100,000 on Tranco and 23.4% to domains outside the top million entirely.
  • Three sites, wifitalents.com, worldmetrics.org and gitnux.org, registered via NameCheap between December 2023 and May 2024 and sharing Cloudflare nameservers and a page template, published 215,128 machine-generated "best " pages between them, versus six real blog posts each.
  • worldmetrics.org and gitnux.org each title their homepage "Facts & Grounding Page" prefixed with their own brand, and describe themselves as a "machine-readable record" of "verified facts," while worldmetrics.org sells custom research from 5,000 euros.
  • Guideflow.com, a product-demo vendor unrelated to any tested category, was the third-most-cited domain overall (194 citations across 96 categories), ahead of Gartner.
  • sonar and sonar-pro shared a byte-identical citation list in 289 of 380 categories (Jaccard 0.898), so the two Perplexity tiers largely reflect one retrieval layer rather than independent measurements; the report covers Perplexity only.

Why it matters

The report's core finding is not that Perplexity cites obscure sites; it is that a measurable share of what a widely used AI search product treats as evidence was built, on a demonstrable timeline and template, specifically to be retrieved by models rather than read by people. The clearest signal is the sites' own self-description: pages titled "Facts & Grounding Page" and meta descriptions offering a "machine-readable record" are addressed to a retrieval system, not to a buyer. That a few such sites can occupy visible positions in an AI's evidence base, generating over 215,000 pages against six real blog posts each, shows how retrieval-grounded answers can be shaped by supply-side content built for the crawler rather than for demand from readers.

Who it affects

Anyone using Perplexity's grounded answers to research software purchases is exposed to citations skewed toward low-traffic, newer domains: 36.5% of the 2,055 cited domains sit outside Tranco's top million, and the median cited domain that does rank sits around position 71,611, far outside the familiar web. It also affects legitimate vendors and review sites competing for visibility, such as Gartner, which the report says guideflow.com's own marketing blog outranked in citation count, and any AI product or research team relying on retrieval-augmented answers built on similar grounding pipelines, since the report notes it did not test ChatGPT, Gemini, Copilot or Google's AI Mode and makes no claim about them.

How to use it

The report and its underlying dataset, including every citation, the Tranco and Wayback lookups, the vendor liveness checks and the scripts that produced each figure, are published at trellner.com under a CC BY 4.0 licence, so the methodology and raw evidence can be independently checked or reused rather than taken on trust. Anyone auditing an AI product's retrieval quality could adapt the same approach: run a fixed category list through a grounded model, capture every returned citation, and rank the resulting domains against an independent popularity measure like Tranco.

How solid is it

The methodology is unusually transparent for this kind of claim: categories were fixed in advance and never revised after seeing results, all 760 API calls returned parseable answers, every one of the 1,502 vendor homepages was independently fetched and cross-checked through a rotating proxy to avoid false dead-site readings, and the full dataset and scripts are released alongside the write-up. The report is also careful about what it does not show: it explicitly flags that sonar and sonar-pro returned identical citation lists in 289 of 380 categories with a 0.898 Jaccard overlap, meaning the two model tiers are not really independent measurements of the same question, and that common ownership of the three sites is inferred from shared Cloudflare nameservers and an identical template rather than proven, since none of the three sites names an operator.

Risks and caveats

The report covers Perplexity's sonar and sonar-pro only, explicitly excludes Google (whose Gemini grounding on OpenRouter would reflect OpenRouter's search plugin rather than Google's own retrieval) and makes no claim about ChatGPT, Gemini, Copilot or Google's AI Mode. It is a single day's snapshot of a retrieval index that changes, drawn from 380 categories the researchers themselves constructed rather than a sample of real buyer queries, with one prompt wording and no repeat sampling; an earlier pilot suggested swapping "best" for "most popular" shifts the product shortlist more than it shifts the citation mix, but this run did not measure that. The report also did not test whether removing the flagged sources would change Perplexity's actual recommendations, and it states plainly that guideflow.com and the three Best List sites may well name reasonable products; the finding is about the composition of the evidence base, not a claim that the recommendations themselves are wrong.

“We do not know who operates them; none of the three names an owner.”

— the report, trellner.com