A business owner asking an AI assistant "what's the best CRM for a five-person team" is doing something that felt, until recently, like a shortcut around bad information: skip the ad-stuffed listicles, skip the sponsored comparison sites, and get a synthesized answer instead. A report published today by Trellner Research suggests that shortcut has already been reverse-engineered.

The researchers put 380 buyer-intent software categories, everything from CRM software to museum collection management software, to Perplexity's two search models through OpenRouter. They kept every citation that came back: 7,534 of them, across 2,055 distinct domains, backing 3,800 recommendation slots. Then they checked where those citations actually pointed.

The sources behind the answer aren't the sources you'd expect

59.8% of those 7,534 citations led to domains ranked worse than the top 100,000 sites on the internet by Tranco's traffic index. 23.4% led to domains that don't appear in the top million at all. The median cited domain ranks 71,611st. For comparison, Wikipedia, the site most people assume dominates AI answers, was cited three times out of 7,534.

That's not one famous source problem. It's the opposite: the answer is built mostly from sites nobody would recognize, and a meaningful chunk of them were built for exactly this purpose. Three connected domains, first registered between December 2023 and May 2024 and sharing the same Cloudflare nameservers, template, and navigation structure, generated 215,128 auto-formatted "best [category] software" pages between them. There are not 215,128 real software categories. Two of the three give their own homepage the HTML title "Facts & Grounding Page." Each carries three named staff bios, nine fictional-sounding names in all, and an "editorial process" page. None of it is written for a person to read. It's written for a retrieval system to fetch, which is precisely what happened.

The most-cited "review" site isn't a review site

The third-largest cited domain in the whole dataset, ahead of Gartner, is guideflow.com, a vendor that sells interactive product demos. It doesn't review software and doesn't compete in the categories it was cited for. Its content-marketing blog was pulled in 194 times across 96 different categories: architecture practice software, RFID software, IVR software, 3D rendering software. Nothing about this is deceptive on Guideflow's end. It's an ordinary company blog. What it reveals is what the retrieval layer does with ordinary content: a vendor's own marketing posts about markets it has nothing to do with became a load-bearing source for "which product should I buy."

Pull the same category from all three manufactured sites and you get three different top-five lists built from the same 2,055-domain pool. For "project estimation software," one site's top pick doesn't appear anywhere in a second site's top five. That's not three independent opinions. It's one supply of machine-generated content, sliced three ways, all feeding the same models the same afternoon.

Citations aren't proof, even when they exist

A second report published the same day, from Haus Research, checked a different failure mode: not where citations point, but whether they say what they're cited for. Researchers asked Perplexity's models 310 factual questions about 210 tech companies and fetched every source attached to a claim involving a specific number. 34.7% of those 1,826 citations pointed at a page that wouldn't open to an ordinary reader or that, once opened, contained none of the figures in the sentence it was supposedly backing. Scored by claim instead of by citation marker, 14.4% of the underlying claims failed outright.

Put the two reports together and the picture is consistent: a citation next to an AI answer signals that a URL exists, not that the URL supports what's being claimed, and not that the URL was written by anyone with a stake in getting it right.

What this means if you're the one asking

This isn't an argument against using AI to research a purchase. It's an argument against treating the citation as the verification step. If you're shortlisting software or a vendor with an AI assistant, the model just did the easy 80% of the work: it read a lot of pages fast. The remaining 20%, checking that a cited source is a real, independent, qualified opinion and not a page built to be fetched, is still yours to do. That's the same discipline we wrote about when a company's "100% human" trust claim turned out to be checkable, and failed the check. A claim, or a citation, is only as good as how easily someone can verify it.

It also cuts the other way for anyone publishing content meant to inform a buying decision, including us. An AI-search strategy that chases citation count by publishing thin, formulaic "best X" pages is optimizing for a metric that two independent reports just showed is already saturated with exactly that kind of page. The sites getting caught in this report weren't rewarded for quality. They were rewarded for volume and machine-readable structure, and now they're the example of what to distrust. If you're evaluating how your own site gets found, built, and cited by AI systems, that distinction between structured-for-machines and actually-useful-to-a-buyer is the one worth building around, not gaming.