跳到正文
RCreddit.com·

[2610.08927] Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

AI 摘要

AI systems have shown rapid progress in scientific discovery with well-defined metrics, but their ability to autonomously perform open-ended discovery is less clear. Researchers investigated this in Station, an open-world environment simulating a scientific ecosystem. By augmenting Station with a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration, agents rediscovered 62.7% of original findings from ICLR papers, significantly outperforming Codex Multiagent-v2 (15.4%) and AI Scientist-v2 (14.4-20.6%). This suggests that a suitable environment can enable AI agents to make meaningful progress in open-ended scientific discovery.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月9日 13:26 UTC

收录当时偏移:UTC+02026年10月9日 15:00 UTC

发布
2026年10月9日 13:26
收录
2026年10月9日 15:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.

来源·reddit.com