返回
Hhackernews·jakobgreenfeld
爆款 · 2.1×31
·5小时前·其他 · 官方 API

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

查看原文
模型发布

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

爆款判定
判定依据
热度约为该来源近期上榜条目中位水平的 2.1 倍
指标对比
197 vs 中位 93.5(20 条基线样本)
检出时间
09/02 18:01
AI 摘要

一项研究发现,当被问及 380 个软件类别中的最佳产品时,两个基于网络的模型引用了 7,534 个 URL。其中,59.8% 的域名在 Tranco 排名中低于第 100,000 位,23.4% 的域名甚至不在前一百万名之内。有三个网站,均在 2023 年 12 月之后创建并似乎受共同控制,它们共同发布了 215,128 个机器生成的“最佳软件”页面,其中两个网站的 HTML 标题为“Facts & Grounding Page”。

We asked two web-grounded models for the best products in 380 software categories and kept every URL they retrieved. Of the 7,534 citations that came back, 59.8% point at domains ranked worse than #100,000 in the Tranco top-1M list and 23.4% at domains that are not in the top million at all. Two of the sites doing the grounding have given their homepage the HTML title “Facts & Grounding Page” — grounding being the retrieval step these models perform — and they and a third site under apparently common control have published 215,128 machine-generated best pages between them; none of the three domains existed before December 2023.

What we ran

On 2 September 2026 we put 380 buyer-intent categories — from “CRM software” to “museum collection management software” — to perplexity/sonar and perplexity/sonar-pro through OpenRouter, one prompt per category per model, 760 calls in all. Each call asked for a ranked top five as JSON, with each product’s official homepage domain. All 760 returned a parseable answer, and both models report the URLs they retrieved, which is why they were chosen. The categories were written before any results were seen and never revised.

That produced 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains. We then looked up every cited domain in the Tranco daily list for 2026-09-01 and in the Wayback Machine, and fetched every one of the 1,502 vendor homepages the models supplied to see whether it still exists.

Google was left out. Grounding a Gemini model on OpenRouter means routing it through OpenRouter’s own web-search plugin, so the citations would describe that plugin rather than Google’s retrieval. Only Perplexity was measured, and nothing here should be read as a claim about any other engine.

Where the citations land

Citations Unranked (outside Tranco 1M) Ranked worse than #100k

perplexity/sonar 3,767 23.4% 59.8%

perplexity/sonar-pro 3,767 23.5% 59.9%

Pooled 7,534 23.4% 59.8%

The median Tranco rank of the 5,768 citations that point at a ranked domain is 71,611. Concentration at the top is unremarkable — the ten most-cited domains take 17.3% of citations — so the story is not that a cartel of famous sites supplies the answers. It is what fills the other four-fifths: 751 of the 2,055 cited domains, 36.5% of them, do not appear in the top million.

Those domains are also newer. The median first Wayback capture is 2020 for the unranked cited domains against 2011 for the ranked ones, and 16.6% of the archived unranked domains were first captured in 2025 or later, against 1.6% of the archived ranked ones.

The ten most-cited domains:

Domain Citations Share Tranco rank

g2.com 291 3.86% 4,027

reddit.com 261 3.46% 105

guideflow.com 194 2.57% 177,039

gartner.com 158 2.10% 1,766

zapier.com 82 1.09% 2,919

wifitalents.com 71 0.94% 105,281

capterra.com 68 0.90% 6,387

linkedin.com 67 0.89% 18

worldmetrics.org 60 0.80% 104,737

gitnux.org 50 0.66% 42,759

Three sites made 215,128 “best software” pages for AI. Perplexity cites them · BuzzRadr