// HACKER NEWS — CYBERSECURITY
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
Across 380 software categories, 59.8% of the sources behind grounded AI recommendations sit outside the 100,000 most-visited websites, and several of the most-cited are sites built to be read by models rather than by people.
We asked two web-grounded models for the best products in 380 software categories and kept every URL they retrieved. Of the 7,534 citations that came back, 59.8% point at domains ranked worse than #100,000 in the Tranco top-1M list and 23.4% at domains that are not in the top million at all. Two of the sites doing the grounding have given their homepage the HTML title “Facts & Grounding Page” — grounding being the retrieval step these models perform — and they and a third site under apparently common control have published 215,128 machine-generated best pages between them; none of the three domains existed before December 2023.
On 2 September 2026 we put 380 buyer-intent categories — from “CRM software” to “museum collection management software” — to perplexity/sonar and perplexity/sonar-pro through OpenRouter, one prompt per category per model, 760 calls in all. Each call asked for a ranked top five as JSON, with each product’s official homepage domain. All 760 returned a parseable answer, and both models report the URLs they retrieved, which is why they were chosen. The categories were written before any results were seen and never revised.
That produced 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains. We then looked up every cited domain in the Tranco daily list for 2026-09-01 and in the Wayback Machine, and fetched every one of the 1,502 vendor homepages the models supplied to see whether it still exists.
Google was left out. Grounding a Gemini model on OpenRouter means routing it through OpenRouter’s own web-search plugin, so the citations would describe that plugin rather than Google’s retrieval. Only Perplexity was measured, and nothing here should be read as a claim about any other engine.
The median Tranco rank of the 5,768 citations that point at a ranked domain is 71,611. Concentration at the top is unremarkable — the ten most-cited domains take 17.3% of citations — so the story is not that a cartel of famous sites supplies the answers. It is what fills the other four-fifths: 751 of the 2,055 cited domains, 36.5% of them, do not appear in the top million.
Those domains are also newer. The median first Wayback capture is 2020 for the unranked cited domains against 2011 for the ranked ones, and 16.6% of the archived unranked domains were first captured in 2025 or later, against 1.6% of the archived ranked ones.
Wikipedia, for comparison, was cited three times in 7,534.
guideflow.com sells interactive product demos. It is not a review site, a directory or a publisher, and it competes in none of the categories we asked about. Its blog was nonetheless cited 194 times across 96 of our 380 categories — a quarter of them — placing it third overall and ahead of Gartner. Each citation is a different URL: 96 distinct guideflow.com blog URLs, one per category, six of them the Estonian-locale copy of a post. Its sitemap lists 3,351 blog URLs, 2,176 of them distinct posts. It supplied the grounding for “3D rendering software”, “IVR software”, “RFID software” and “architecture practice software” alike.
Nothing here is deceptive. Guideflow publishes a large content-marketing blog, as thousands of companies do. The measurement is about what the retrieval layer does with it: a vendor’s own listicles about markets it does not operate in became the third-largest evidence base for a question about which product to buy.