返回
RCreddit.com
18
·14小时前·RSS
暂不在当前实时榜单

Which LLM is actually best at pentesting? benchmark to find out

查看原文
模型发布

热度趋势

新上榜
最近 24 小时与此前 24 小时对比 · 7 天曲线

百分比基于当前可用热度信号,而非评论数或独立用户人数。

AI 摘要

A new benchmark has been developed to assess which Large Language Models (LLMs) are most effective for penetration testing. The creator found existing benchmarks, like CyberGym, to be insufficient as they focus on exploit generation for known vulnerabilities rather than real-world pentesting scenarios. This new benchmark provides LLMs with live infrastructure to attack, aiming to offer a more accurate comparison of their capabilities in a pentest-oriented environment. The project, initially personal, is now shared for broader use.