RCreddit.com·
暂不在当前实时榜单
Turns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.
Many current science-based LLM benchmarks have been found to contain flaws in their answers. When these benchmarks were corrected, the scores of the LLMs being evaluated rose significantly. This discovery suggests that the true capabilities of LLMs in scientific domains might be underestimated due to issues within the benchmarks themselves. This information comes from a paper posted on Arxiv, as discussed in a dev_community on reddit.com.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月16日 10:00 UTC
- 收录
- 2026年9月16日 10:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →