I gave Claude models a web_search tool and 40 questions. Opus decided right 79 of 80 times. Sonnet 5 with no system prompt told me Harald V is still king.
A developer tested Claude models with a web_search tool on 40 questions, half of which were about events after the models' training cutoff. The goal was to see if models would use the search tool when the answer depended on recent information. Opus correctly decided to use the search tool 79 out of 80 times, while Sonnet 5, without a system prompt, incorrectly stated that King Harald V was still alive. The developer spent $6.53 on 2000 calls across 16 models via OpenRouter.
This report uniquely details a developer's direct testing of Claude models' web search tool integration, providing specific performance metrics for Opus and Sonnet 5 unlike general capability claims.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 2, 2026, 14:00 UTC
- Ingested
- Oct 2, 2026, 14:00
- Source type
- Dev community
Full text isn't available here.
Read at source →