RCreddit.com·

OpenAI say they already have an automated AI research intern and expect a full AI researcher by March 2028 which could lead to RSI. Is this legit or just IPO hype before the bubble pops? https://openai.com/index/research-acceleration-view-inside-openai/

AI summary

OpenAI claims to have an automated AI research intern and anticipates a full AI researcher by March 2028, potentially leading to Recursive Self-Improvement (RSI). This is supported by Astra's performance in ARC-AGI-3, where it scored 62.7% with a standard harness and 99.9% with an OpenAI-specific harness, using fewer actions than most humans on 96% of solved levels. The community questions if this is legitimate progress or merely IPO hype.

Time & source
Published
Sep 6, 2026, 18:52
Source type
Dev community
Tier
Community
Source status
Healthy
Tier is a per-source editorial setting, not a per-item score.

Times shown in UTC

More details
First seenSep 7, 2026, 04:00Time zoneUTC · UTC+0
Article

For me, the answer is obvious, but I'm interested to hear the opinion of others to understand if my biases are clouding my view.

The most persuasive evidence for me is the improvements in the models. For Astra, this is best demonstrated in agentic behaviour, which is critical for having an automated AI researcher, as agentic behaviour means the model can actually do things. Use a computer. Write code. Run code. Use tools. Search. Carry out experiments. Look at the results. Change what it’s doing. Try again. Keep working on a problem.

And if the brain keeps improving at the same time as the hands improve, it seems completely plausible to me that the job of an AI researcher will eventually become automatable.

For those who don't share this view, what does an automated AI researcher fundamentally need to be able to do that a few more iterative model improvements won’t plausibly give it?

Even if you think that the jumps required are large, it's important to note that models are making significant improvements.

For example, the jump on SpatialBench for Astra is ridiculous. The benchmark consists of 25 visual path-tracing problems and 25 3D mental-rotation problems, which are exactly the sort of things previous multimodal models were embarrassingly poor at. In fact, GPT-5.5 scored only around 13%. But Astra scored around 67% without tools and over 90% with Code Interpreter, which is above the reported human baseline of 80%.

Then you have the ARC-AGI-3 results, where models are required to enter unfamiliar environments, work out the rules, build representations of what’s going on, plan, use tools and create software to help solve the problem. GPT-5.5 previously scored just 0.43%. Astra scored 62.7% under the standard provider-neutral harness, while an OpenAI-specific harness that preserved its reasoning state reached 99.9%. Remarkably, Astra used fewer actions than the median human on 96% of the levels it solved.

So what am I missing? What do sceptics think will stop this progression before we get to an automated AI researcher? And if we get that far, what stops us from achieving RSI? And if we get that far, what does a post-RSI world look like?