The overlooked GPT-5.6 metric: OpenAI’s internal AI usage jumped 22x in six months
Heat trend
The percentage is based on available heat signal, not comment count or independent people.
This covers a coding tool or code-capability update — useful for developers assessing workflow changes and reusable value.
Everyone is focused on the benchmark scores, but one number in OpenAI’s GPT-5.6 release caught my attention much more. OpenAI says its internal agentic token usage increased roughly 22x over the last six months, while the share of research compute going toward internal coding inference grew 100x. That feels more significant than another model gaining a few points on a benchmark. The company building frontier models is increasingly using those same models inside the process of building the next generation. I don’t think this means “recursive self-improvement” in the dramatic sci-fi sense. Humans are still setting the objectives, running the infrastructure and deciding what gets trained. But it does look like the R&D loop is starting to compress: better models help researchers work faster, which helps produce better models, which can then accelerate the next cycle. If that trend continues, the important question may stop being how much better is the next model? and become how much faster can an AI lab improve once AI is deeply embedded in its own research process? Curious where people here draw the line between ordinary AI-assisted engineering and something qualitatively different.