On Measurement, Scaling, and the Illusion That Everything Can Be Optimized
人工智能模型正快速发展,能力和多功能性不断增强,甚至在数月前看似遥不可及的任务上取得成功,这使得人们再次声称已达到通用人工智能(AGI)的门槛。然而,一个简单的数学问题“x + 5 = 3”揭示了现有评估方法的局限性。AI的决定性失败并非基准所能衡量,而是基准甚至无法提出的问题。这表明需要重新审视问题空间,并引入具体的数学原语作为新的评估候选。
- 发布
- 2026年9月7日 10:24
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
时间以 UTC 显示
更多信息
AI is progressing fast. Models are becoming more capable, more versatile, and are succeeding at tasks that seemed out of reach just months earlier. With every new generation, the same story returns. This time, we are supposed to have crossed a threshold, reached AGI.
Ever more complex mathematical problems are being solved. Task horizons are getting longer. Scientific performance is improving. Token consumption is exploding. Gradually, mathematics, science, and the ability to execute long tasks are becoming benchmarks for intelligence itself.
We have learned to measure, then to optimize what we measure. The danger begins when we mistake this optimization for a deeper understanding of the problem. What can be measured eventually comes to stand in for a definition. Mechanical intelligence takes the place of discernment.
The Limits of Brute Force
By brute force, I mean optimization in the broad sense, using more compute, more data, more context, more inference time, more search, more verification, more tools, or more attempts to extract additional performance from a system. Scaling is now the most visible expression of this logic.
This logic can go very far and continually push back the operational limits of models. But further optimizing a system can considerably extend what it is able to do without changing the space in which that progress takes place.
This distinction between gaining performance and changing the framework is what so often disappears from the narrative of progress.
Pushing Back a Limit Is Not Changing the Framework
Imagine a mathematical system that only knows strictly positive numbers. You can improve its speed, precision, and algorithms as much as you like. As long as you evaluate it on problems whose solutions are positive, its performance can become perfect.
Now take the equation x + 5 = 3. No positive number can solve it. Giving the system a thousand times more computing power changes nothing. The solution, x = -2, is not difficult to find. It simply does not exist within the space of admissible solutions.
Negative numbers do not make the system better within its framework. They change the framework by making representable a solution that could not be represented before.
That is what a paradigm shift is.
The Failure the Benchmark Cannot See
A system limited to positive numbers can score 100% on every benchmark built within the space of positive numbers. If those benchmarks define its capabilities, we will eventually declare the system complete and mistake the limits observed within that space for its absolute limits.
Yet simply posing x + 5 = 3 reveals what every previous evaluation was incapable of seeing. The decisive failure is therefore not the one the benchmark measures. It is the one the benchmark cannot even formulate.
That is exactly the problem when mathematics, science, success rates, or task horizons become measures of intelligence itself. We optimize models on what we know how to formalize, then evaluate them on what we know how to measure. Worse, we end up defining intelligence through those same measures.
Every saturated benchmark then produces two results, a success in optimization for the machine and the revelation of a limit in our measurement.
When a benchmark stops discriminating, we simply move the target. We build harder problems, longer horizons, and more demanding criteria. Scores drop, the model optimizes again, and the next saturation point is presented as a new frontier of intelligence, when it may simply be a new horizon of optimization, increasingly expensive in energy and capital.
We know how to measure the extension of performance. We know far less about how to measure the extension of the framework. A benchmark can show that the system solves more problems. It cannot show that it has discovered negative numbers.
What we are encountering, then, is not necessarily a performance wall. It is a conceptual wall. Scaling can continue to produce considerable gains without allowing us to determine whether we are simply exploring the same space more effectively or whether we have genuinely changed the framework. The problem is not that optimization has stopped working. It is that optimization alone is no longer enough to demonstrate what we claim it demonstrates.
What Resists Optimization, Our Negative Numbers
The black box is a first example. We know how to train and optimize systems whose internal workings we cannot precisely reconstruct. We observe that they perform better without always knowing whether that improvement comes from a new representation, a better use of what was already there, or a more effective way of working around their limitations.
We are extremely effective at optimizing what we know how to formalize. The problem begins when we forget the assumptions that made this formalization possible and end up mistaking them for properties of reality itself.
The danger is not merely that we remain trapped within positive numbers. It is that we forget that "positive numbers only" was a starting assumption.