Z.ai’s GLM-5.3 model was trained with post-training scaling improvements aimed at vulnerability detection, achieving 84.5% on the CyberGym benchmark for reproducing known flaws in real codebases. Through a vulnerability-discovery program run with Chinese security teams, the model identified 2,436 vulnerabilities across 269 open-source projects, with a mean dormancy of 26.6 years before discovery. The model still trailed closed frontier competitors by 23.6 points on the more demanding ExploitBench task, suggesting its bug-finding capability is currently outpacing its exploit-development capability.
