RCreddit.com·

OpenAI say they already have an automated AI research intern and expect a full AI researcher by March 2028 which could lead to RSI. Is this legit or just IPO hype before the bubble pops? https://openai.com/index/research-acceleration-view-inside-openai/

AI 摘要

OpenAI声称已拥有一个自动化AI研究实习生,并预计到2028年3月将实现一个完整的AI研究员,这可能导致递归自我改进(RSI)。这一说法得到了Astra在ARC-AGI-3测试中表现的支持,该模型在标准测试中得分62.7%,而在OpenAI特定测试中得分高达99.9%,并且在96%的已解决关卡中,其行动次数少于人类中位数。社区对此进展的真实性表示疑问,认为这可能只是首次公开募股前的炒作。

时间与来源
发布
2026年9月6日 18:52
来源类型
开发者社区
档位
社区
信源状态
正常
档位是按信源手工设定的编辑判断,不是逐条打分。

时间以 UTC 显示

更多信息
首次发现2026年9月7日 04:00时区UTC · UTC+0
正文

For me, the answer is obvious, but I'm interested to hear the opinion of others to understand if my biases are clouding my view.

The most persuasive evidence for me is the improvements in the models. For Astra, this is best demonstrated in agentic behaviour, which is critical for having an automated AI researcher, as agentic behaviour means the model can actually do things. Use a computer. Write code. Run code. Use tools. Search. Carry out experiments. Look at the results. Change what it’s doing. Try again. Keep working on a problem.

And if the brain keeps improving at the same time as the hands improve, it seems completely plausible to me that the job of an AI researcher will eventually become automatable.

For those who don't share this view, what does an automated AI researcher fundamentally need to be able to do that a few more iterative model improvements won’t plausibly give it?

Even if you think that the jumps required are large, it's important to note that models are making significant improvements.

For example, the jump on SpatialBench for Astra is ridiculous. The benchmark consists of 25 visual path-tracing problems and 25 3D mental-rotation problems, which are exactly the sort of things previous multimodal models were embarrassingly poor at. In fact, GPT-5.5 scored only around 13%. But Astra scored around 67% without tools and over 90% with Code Interpreter, which is above the reported human baseline of 80%.

Then you have the ARC-AGI-3 results, where models are required to enter unfamiliar environments, work out the rules, build representations of what’s going on, plan, use tools and create software to help solve the problem. GPT-5.5 previously scored just 0.43%. Astra scored 62.7% under the standard provider-neutral harness, while an OpenAI-specific harness that preserved its reasoning state reached 99.9%. Remarkably, Astra used fewer actions than the median human on 96% of the levels it solved.

So what am I missing? What do sceptics think will stop this progression before we get to an automated AI researcher? And if we get that far, what stops us from achieving RSI? And if we get that far, what does a post-RSI world look like?