Skip to content
RCreddit.com·

On Measurement, Scaling, and the Illusion That Everything Can Be Optimized

AI summary

The rapid progress of AI models, which are becoming more capable and versatile, has led to claims of reaching Artificial General Intelligence (AGI). However, a critical flaw in current AI evaluation methods is highlighted by simple problems like "x + 5 = 3." This suggests that benchmarks fail to identify decisive failures because they cannot even formulate the right questions, indicating a need to redefine the problem space with new mathematical primitives.

Time & source
Published
Sep 7, 2026, 10:24
Source type
Dev community
Tier
Community
Source status
Healthy
Tier is a per-source editorial setting, not a per-item score.

Times shown in UTC

More details
First seenSep 7, 2026, 14:00Time zoneUTC · UTC+0
Article

AI is progressing fast. Models are becoming more capable, more versatile, and are succeeding at tasks that seemed out of reach just months earlier. With every new generation, the same story returns. This time, we are supposed to have crossed a threshold, reached AGI.

Ever more complex mathematical problems are being solved. Task horizons are getting longer. Scientific performance is improving. Token consumption is exploding. Gradually, mathematics, science, and the ability to execute long tasks are becoming benchmarks for intelligence itself.

We have learned to measure, then to optimize what we measure. The danger begins when we mistake this optimization for a deeper understanding of the problem. What can be measured eventually comes to stand in for a definition. Mechanical intelligence takes the place of discernment.

The Limits of Brute Force

By brute force, I mean optimization in the broad sense, using more compute, more data, more context, more inference time, more search, more verification, more tools, or more attempts to extract additional performance from a system. Scaling is now the most visible expression of this logic.

This logic can go very far and continually push back the operational limits of models. But further optimizing a system can considerably extend what it is able to do without changing the space in which that progress takes place.

This distinction between gaining performance and changing the framework is what so often disappears from the narrative of progress.

Pushing Back a Limit Is Not Changing the Framework

Imagine a mathematical system that only knows strictly positive numbers. You can improve its speed, precision, and algorithms as much as you like. As long as you evaluate it on problems whose solutions are positive, its performance can become perfect.

Now take the equation x + 5 = 3. No positive number can solve it. Giving the system a thousand times more computing power changes nothing. The solution, x = -2, is not difficult to find. It simply does not exist within the space of admissible solutions.

Negative numbers do not make the system better within its framework. They change the framework by making representable a solution that could not be represented before.

That is what a paradigm shift is.

The Failure the Benchmark Cannot See

A system limited to positive numbers can score 100% on every benchmark built within the space of positive numbers. If those benchmarks define its capabilities, we will eventually declare the system complete and mistake the limits observed within that space for its absolute limits.

Yet simply posing x + 5 = 3 reveals what every previous evaluation was incapable of seeing. The decisive failure is therefore not the one the benchmark measures. It is the one the benchmark cannot even formulate.

That is exactly the problem when mathematics, science, success rates, or task horizons become measures of intelligence itself. We optimize models on what we know how to formalize, then evaluate them on what we know how to measure. Worse, we end up defining intelligence through those same measures.

Every saturated benchmark then produces two results, a success in optimization for the machine and the revelation of a limit in our measurement.

When a benchmark stops discriminating, we simply move the target. We build harder problems, longer horizons, and more demanding criteria. Scores drop, the model optimizes again, and the next saturation point is presented as a new frontier of intelligence, when it may simply be a new horizon of optimization, increasingly expensive in energy and capital.

We know how to measure the extension of performance. We know far less about how to measure the extension of the framework. A benchmark can show that the system solves more problems. It cannot show that it has discovered negative numbers.

What we are encountering, then, is not necessarily a performance wall. It is a conceptual wall. Scaling can continue to produce considerable gains without allowing us to determine whether we are simply exploring the same space more effectively or whether we have genuinely changed the framework. The problem is not that optimization has stopped working. It is that optimization alone is no longer enough to demonstrate what we claim it demonstrates.

What Resists Optimization, Our Negative Numbers

The black box is a first example. We know how to train and optimize systems whose internal workings we cannot precisely reconstruct. We observe that they perform better without always knowing whether that improvement comes from a new representation, a better use of what was already there, or a more effective way of working around their limitations.

We are extremely effective at optimizing what we know how to formalize. The problem begins when we forget the assumptions that made this formalization possible and end up mistaking them for properties of reality itself.

The danger is not merely that we remain trapped within positive numbers. It is that we forget that "positive numbers only" was a starting assumption.

Source·reddit.com·Full text via RSS