跳到正文
RCreddit.com·
暂不在当前实时榜单

I gave Claude models a web_search tool and 40 questions. Opus decided right 79 of 80 times. Sonnet 5 with no system prompt told me Harald V is still king.

AI 摘要

A developer tested Claude models with a web_search tool on 40 questions, half of which were about events after the models' training cutoff. The goal was to see if models would use the search tool when the answer depended on recent information. Opus correctly decided to use the search tool 79 out of 80 times, while Sonnet 5, without a system prompt, incorrectly stated that King Harald V was still alive. The developer spent $6.53 on 2000 calls across 16 models via OpenRouter.

为什么是这条

This report uniquely details a developer's direct testing of Claude models' web search tool integration, providing specific performance metrics for Opus and Sonnet 5 unlike general capability claims.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月2日 14:00 UTC

收录
2026年10月2日 14:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com