返回
RCreddit.com
21
·9小时前·开发者社区 · RSS

Z.ai ships GLM-5.3, holds open weights for cyber safety review

查看原文
OpenAI模型发布

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

OpenAI 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。

AI 摘要

Z.ai 发布了 GLM-5.3,声称其性能较 GLM-5.2 有显著提升,这些提升主要通过扩展的后期训练而非重新预训练实现。该公司报告称,其内部基准测试取得了显著进展:Terminal-Bench 3.0 从 4.6 升至 28.3,DeepSWE v1.1 从 46.2 升至 66.9,CyberGym 达到 84.5%。…

Post-training alone did the heavy lifting on Z.ai's latest release, and that is the part worth pausing on. In the [z.ai launch post]( https://z.ai/blog/glm-5.3 ), the lab said GLM-5.3 runs on the same mixture-of-experts base as GLM-5.2 and every reported gain came from extended post-training rather than a fresh pretrain. The scoreboard the company is putting out is aggressive: Terminal-Bench 3.0 climbs from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and CyberGym reaches 84.5%, edging Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

The security numbers are what make this launch different from the routine coding-benchmark press release. Z.ai says the model surfaced 2,436 vulnerabilities across 269 open-source projects during evaluation, with 1,097 rated critical or high severity, and reports finding critical bugs in Linux, WebKit, and FreeBSD. The lab also says the model began reasoning across multiple stages of exploitation and forming coherent plans for complete exploitation chains, a capability it did not set out to train for. That admission is why weights are being held back roughly two weeks for safety evaluation and hardening, according to reporting from [SiliconANGLE]( https://siliconangle.com/2026/08/14/z-ai-debuts-glm-5-3-long-horizon-coding-cybersecurity-upgrades/) and [The Agent Report]( https://the-agent-report.com/2026/08/glm-5-3-zai-post-training-coding-cyber/ ). For a lab whose open-weight releases are much of the reason its models get attention, that is not a small choice.

Some caution is warranted on the specifics. Every score above comes from Z.ai's own report on its own benchmark mix, so independent reruns have not landed yet, and coverage notes GLM-5.3 still trails Fable 5 and GPT-5.6 Sol badly on ExploitBench and ExploitGym. The launch post does not describe what criteria decide whether the weights actually ship in two weeks or who signs off, and nothing in the reporting pins down whether the 2,436 disclosed bugs were coordinated with the affected maintainers first.

Z.ai ships GLM-5.3, holds open weights for cyber safety review · BuzzRadr