I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
An experiment tested over 20 methods to make a cheap coding model, Haiku, perform like a more expensive one, Sonnet. The study involved pre-registered experiments on real repository commits, with protocols committed to Git before execution. A key finding was that running the agent's change and reporting facts, such as "if this line became pass, all tests would still pass," was more effective than simply giving advice, achieving 35/42 successful outcomes compared to 32/42, and reducing formatting regressions from 10 to 0.
This report details the first pre-registered experiments comparing methods to improve a cheaper coding model's performance, unlike many anecdotal accounts.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 6, 2026, 05:00 UTC
- Ingested
- Oct 6, 2026, 05:00
- Source type
- Dev community
Full text isn't available here.
Read at source →