返回
Hhackernews·jeudesprits
49
·10小时前·官方发布 · 官方 API

GLM-5.3 is now open-weight

查看原文
官方公告Hugging Face

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

- Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam.

- Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.

Benchmark

Benchmark GLM-5.3 GLM-5.2 Kimi K3 DeepSeek-V4 Pro-0813 Qwen3.8-Max Opus 4.8 Fable 5 (w/ fallback) GPT-5.6 Sol

Terminal Bench 2.1 88.2 81.0 88.3 87.9 86.6 85.0 88.0 88.8

Terminal Bench 3.0 28.3 4.6 17.4

21.1 33.7 34.6

DeepSWE (v1.1) 66.9 46.2 67.5 62.7 56.6 58.0 69.7 72.7

NL2Repo 58.0 48.9 58.0 61.1 55.9 69.7

ProgramBench (Almost Solved) 19.0 9.5 17.5

10.5 15.5 33.0 23.0

FrontierSWE 78.1 67.5

66.5 88.2

SWE-Marathon (v1.1) 42.5 19.4 48.1

48.8 33.1 42.5

PostTrainBench 39.8 31.7 32.0

32.9 41.8 36.2

CyberGym 84.5 77.2 80.0 83.3 78.5 78.1 83.8 83.6

ExploitGym (2h / 6h) 105 / 130 29 / 39 36 / 70

14 / 26 80 / 120 181 / 247 216 / 293

ExploitBench 54.4 24.4 32.2

28.8 40.0 78.0 76.5

Toolathlon Verified 73.0 59.9 76.5 74.1 72.5 76.2 74.7 74.9