One AI just scored 1753 on a test where 'human expert' is 1000. Here's why I don't fully trust that number
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
OpenAI 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。
reddit.com上的一篇帖子对AI在测试中获得1753分(人类专家为1000分)的意义提出了质疑。发帖人对这类旨在病毒式传播的新基准比较表示怀疑,并列举了Grok 4.6、Grok 4.5、GPT-5.6和Fable 5等例子。核心问题在于这些模型是否能真正产出高质量、可交付的工作,或者仅仅是看起来很完善。…
have seeing this kind of new benchmark comparison (Grok 4.6, Grok 4.5, GPT-5.6, Fable 5) and it's the kind of number that's built to go viral.
The test is basically "give the AI real work to do, write docs, spreadsheets, decks and score it against a baseline of what a human expert produces." Human expert sits around 1000. Grok 4.6 scored 1753. On paper that reads as "AI now beats human professionals by a mile."
Except here's the catch no big AI companies is mentioning, this test doesn't grade against a right/wrong answer key. It works by having one AI's output compared side by side against another AI's output, and a grader just picks which one "looks better." There's no ground truth, just a preference vote between two pieces of work.
That setup should ring a bell if you've followed AI chatbot leaderboards before, because those work the same way and people already don't fully trust them. Preference-based grading tends to reward stuff that looks polished, confident, and well-formatted, not necessarily stuff that's actually correct or useful.
A slide deck with clean formatting and a very confident tone can beat a messier but more accurate one, even if the accurate one is the better piece of work. So a 1753 might genuinely mean "produces very professional-looking output." It doesn't automatically mean "does the job better than a human expert."
Doesn't mean the number is fake or the model isn't impressive, it clearly is. Just means I'd treat "AI beats human experts" headlines from this kind of test with a big grain of salt until it's checked against something with an actual correct answer, not just a beauty contest between two AIs.
Curious if anyone's actually used one of these models for real deliverable-style work and can say whether the output holds up, or if it's just really good at looking finished.