Terence Tao says AI labs’ race to beat math benchmarks is starting to hurt the field. He wants them to compete on new insights instead.
Terence Tao expressed concern that AI labs' focus on beating mathematical benchmarks, such as lowering the bound on gaps between primes from 246 to 240, is detrimental to the field. He argues that the specific number is less important than the new insights and tools developed during the process, citing examples like Zhang's work on equidistribution and Maynard's sieve. Tao proposes that AI labs should instead compete on generating genuinely new mathematical insights to foster more meaningful advancements.
- Published
- Sep 6, 2026, 22:15
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Times shown in UTC
More details
On Aug. 31, Stadlmann posted a preprint lowering the bound on gaps between primes from 246 to 240. Within days, several AI labs were posting their own improvements on social media. Tao responded with an eight-post thread explaining why he finds this worrying.
His point is that the number itself was never what mattered most. Going from 246 to 240, or even from 70 million to 246, does little for the rest of mathematics on its own. What mattered was everything developed along the way: Zhang (who was working in a sandwich shop at the time) bringing neglected work on equidistribution back into use, Maynard developing a sieve that became a standard tool, and Polymath8 showing what open collaboration could accomplish. Those advances had applications well beyond the original problem.
Tao imagines how things might have played out if today’s AI labs had been around in 2005, when GPY published their near-miss. The labs pour millions into compute, push the bound into the low hundreds, then move on once progress slows. Mathematicians decide the problem has been picked over and look elsewhere. Nobody writes a proper paper or turns the arguments into something people can learn from. An idea like Maynard’s sieve ends up buried in hundreds of pages of AI output that nobody reads. Zhang never gets his moment; Maynard leaves the field.
In that scenario, the bound improves faster, but mathematics loses out.
It’s a Goodhart’s law problem: the number becomes the target, and the reasons anyone cared about it get lost. Tao doesn’t think AI has to work this way. He points to the Erdős problem example as a case where collaboration with AI helped advance understanding. His objection is to labs bypassing experts and peer review to rush out a better number. He argues that an approach that was relatively harmless in 2025, when models couldn’t solve whole problems unassisted, has started doing more harm than good in 2026.
His proposal is to change what the labs compete over: who can announce a genuinely new mathematical insight first?