RCreddit.com
18
·14小时前·RSS
暂不在当前实时榜单
Which LLM is actually best at pentesting? benchmark to find out
模型发布
热度趋势
新上榜
百分比基于当前可用热度信号,而非评论数或独立用户人数。
A new benchmark has been developed to assess which Large Language Models (LLMs) are most effective for penetration testing. The creator found existing benchmarks, like CyberGym, to be insufficient as they focus on exploit generation for known vulnerabilities rather than real-world pentesting scenarios. This new benchmark provides LLMs with live infrastructure to attack, aiming to offer a more accurate comparison of their capabilities in a pentest-oriented environment. The project, initially personal, is now shared for broader use.