
Ask a web-grounded assistant which software to buy and you get a confident ranked answer with sources attached. Trellner Research, an independent group that studies AI answer engines and takes no payment from the companies it covers, ran that measurement on 2 September 2026, the day before this article, and published it the same day, along with the full dataset and the scripts behind every figure under a CC BY 4.0 licence. It is one day's snapshot of a retrieval index that changes, which Trellner says plainly and which is worth holding onto while reading the numbers.
The short version: the evidence base is mostly not the web you would have searched yourself.
What was actually measured#
Trellner put 380 buyer-intent categories, from CRM software to museum collection management software, to Perplexity's sonar and sonar-pro models through OpenRouter. One prompt per category per model, 760 calls, every one returning a parseable ranked top five with each product's official homepage. That produced 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains.
Those two models were chosen because both report the URLs they retrieved. Google was deliberately excluded: grounding a Gemini model on OpenRouter routes it through OpenRouter's own search plugin, so the citations would describe the plugin rather than Google's retrieval. Trellner states plainly that nothing in the report should be read as a claim about any other engine, and that limit is worth carrying rather than dropping.
Where the citations land#
Of the 7,534 citations, 59.8 percent point at domains ranked worse than 100,000th on the Tranco list of most-visited sites. 23.4 percent point at domains outside the top million entirely. Among the citations that do hit a ranked domain, the median rank is 71,611.
This is not a story about a cartel of famous sites supplying the answers. The ten most-cited domains take only 17.3 percent of citations between them, which is unremarkable concentration. It is a story about the other four-fifths. Of the 2,055 domains cited, 751 do not appear in the top million at all.
Those domains are also younger. Trellner reports a median first Wayback capture of 2020 for the unranked cited domains against 2011 for the ranked ones, and 16.6 percent of the archived unranked domains were first captured in 2025 or later, against 1.6 percent of the ranked.
One number gives the shape of it. Wikipedia was cited three times in 7,534.
A demo vendor's blog was cited more often than Gartner#
The third most-cited domain in the run was guideflow.com, with 194 citations, ahead of Gartner's 158. Guideflow sells interactive product demos. Trellner notes it is not a review site, a directory or a publisher, and competes in none of the categories the models were asked about.
Its blog supplied grounding across 96 of the 380 categories, a quarter of them, on a different URL each time. Trellner counted 96 distinct Guideflow blog URLs in the citation set, and a sitemap listing 3,351 blog URLs. The same blog grounded answers about 3D rendering software, IVR software, RFID software and architecture practice software alike.
Trellner is explicit that this is not an accusation: "Nothing here is deceptive." Guideflow publishes a large content-marketing blog, as thousands of companies do. The measurement is about what the retrieval layer did with it. A vendor's own listicles about markets it does not operate in became the third-largest evidence base for a question about what to buy.
Pages that name the machine they are written for#
Three further sites in and just below the top ten were wifitalents.com, worldmetrics.org and gitnux.org, together accounting for 181 citations across 41 categories.
Trellner takes the three to be one operation, and is careful to say it cannot prove that. It calls a shared nameserver pair "strong circumstantial evidence of a common Cloudflare account rather than proof of ownership", and states flatly: "We do not know who operates them; none of the three names an owner." Nothing that follows is a finding about who owns anything, and it should not be read as one.
What Trellner did observe is a set of matching properties. It reports that all three were registered through the same registrar between December 2023 and May 2024, delegate DNS to the same pair of Cloudflare nameservers, run the same page template with the same navigation, and each keep a blog of exactly six posts, all eighteen of which are about the other brands in the set.
The scale is the finding. Their sitemaps list roughly 105,000 URLs each, of which about 70,000 apiece are generated best-of pages for a software category: 215,128 buying guides across the three brands, against six blog posts each. As Trellner puts it, there are not 215,128 software categories.
Then the detail that is hard to unsee. Fetched on 2 September 2026, two of the three returned an HTML title ending "Facts & Grounding Page", the brand name sitting in front of it, with a matching meta description offering "Company, legal, methodology, and compliance details in one machine-readable record."
Grounding is not a word buyers use. It is the name of the retrieval step in which a system fetches documents to condition an answer on. A machine-readable record of verified facts about oneself is not a service to a human reader either. These pages announce, in their own titles, who they are addressed to.
The same question, three different winners#
Trellner fetched the same category page, project estimation software, from all three brands. Each states its ranking in JSON-LD, so it can be read without interpretation.
Worldmetrics ranked Float, Scoro and Teamwork.com in its top three. WifiTalents ranked the same three, then diverged. Gitnux led with Saviom, which does not appear in Worldmetrics' top five at all.
Each page credits three named staff, nine distinct people across the three sites for one question, and each announces an editorial process, one labelling its result "AI-verified · Expert reviewed". All three carry an unrendered template variable in the byline, reading "Within the next 26 days" on two of them and "Within the next 40 days" on the third.
Where the recommendations pointed#
The models also supplied 1,502 vendor homepages. Trellner checked each twice, directly and through a rotating proxy, counting a site as reachable if either attempt succeeded so that a host blocking one address is not recorded as a dead company.
Most were fine. Seventeen, 1.1 percent, were gone or unreachable, and Trellner is careful to note that four of those are large sites that are plainly alive and simply never answered an automated request. Another 92 redirect to a different registrable domain, mostly ordinary acquisitions and rebrands.
Two failures were of a different kind, and in both the two tiers disagreed with each other. Asked for research data management platforms, one tier gave the real Dryad repository and the other gave a domain that redirects to an Indonesian online-gambling portal. Asked for data quality tools, one tier gave Monte Carlo's actual domain and the other gave a domain belonging to the Monaco hotel and casino group.
What the report does not claim#
The limitations section is unusually good, and reading it changes what the study means.
The two models are not two measurements. They returned byte-identical citation lists in 289 of 380 categories and their URL sets overlap at a Jaccard of 0.898, so Trellner treats the Perplexity tiers as one retrieval stack sampled twice rather than as independent systems that happen to agree.
The 380 categories are Trellner's own construction, not a sample of what buyers really ask, and a list weighted toward niche verticals will surface more long-tail sources than common queries would. It is one prompt wording, one run per category, no repeat sampling, and one day's snapshot of an index that changes. Every fetch went through a datacentre proxy under a named research user-agent, so what these sites returned to Trellner is not necessarily what they return to a retrieval crawler or to a browser.
Trellner also says Tranco rank "is a popularity measure, not a quality measure, and a low rank is not an accusation". And the sentence that keeps the whole thing honest: "We have not shown that any of this changes the answers." Nobody tested whether removing these sources produces different recommendations, and the products named may well be reasonable ones. What was measured is which documents the evidence base is made of.
What to do with this#
The useful move is available in the same place the problem is. Both tiers Trellner tested report the URLs they retrieved, and so do most grounded assistants. Ask for the sources behind a recommendation and open two or three of them before you act on it. If the page ranking your shortlist is a generated best-of on a domain you have never heard of, you have learned something the ranked list itself could not tell you.
The test that falls out of this report is a good one to keep: is this page addressed to a reader or to a retriever? A page written for a person names its author, dates its claims and links to where each number came from. Our own cloud GPU provider comparison is in exactly the category the report is about, and the only defensible answer is the same one demanded of everyone else, that every figure traces to the vendor's own published page.
For anyone buying on behalf of a team, this belongs next to the other risks that do not appear on an invoice, alongside what your coding assistant sends home and the business exposure in open versus closed models. And it is a reminder that a fluent answer is not a checked one, which is the same lesson prompt injection teaches from the other direction.
None of this makes assistants useless for software research. It makes them a starting point with a bibliography, and the bibliography is the part worth reading.



Discussion