Zhipu says GLM-5.3 beats Anthropic, OpenAI at finding bugs

Zhipu says GLM-5.3 beats Anthropic, OpenAI at finding bugs

Chinese company Zhipu launched a new AI model called GLM-5.3 last week, and its announcement includes benchmark data claiming the model beats Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol on CyberGym, a benchmark that tests a model's ability to solve real-world cybersecurity challenges. No specific scores or the margin by which GLM-5.3 is said to lead are given. In its announcement, Zhipu wrote: "As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain." The company added that the model "did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains." Zhipu said it worked with unnamed Chinese companies to run GLM-5.3 against real-world codebases, and the model found 2,436 vulnerabilities across 269 projects, including 1,097 rated medium to high severity. The findings spanned system kernels, operating systems, browser engines, open-source infrastructure, web applications and network protocols. Zhipu's announcement states that many of the flaws had gone unnoticed for years or decades, with the oldest dating back roughly 40 years. The Register does not report whether any of the 2,436 vulnerabilities have since been disclosed, patched or assigned CVEs, nor does it name the companies or codebases involved. The same announcement acknowledges that GLM-5.3 performed worse than western models on other, unnamed security and coding benchmarks. Register APAC editor Simon Sharwood argues that despite that weaker showing elsewhere, a Chinese model being a highly capable bug finder signals that China is not far behind in the ability to find flaws in rival software, and that it developed that capability quickly after the debut of Anthropic's Mythos. He concludes that any advantage the US felt it had as the home of Anthropic has dissipated as a result.

Key facts

  • Zhipu launched GLM-5.3 last week and says it beats Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol on the CyberGym real-world vulnerability-finding benchmark, with no exact scores or margin given.
  • Testing GLM-5.3 with Chinese companies on real-world codebases turned up 2,436 vulnerabilities across 269 projects, including 1,097 rated medium to high severity.
  • Some of the flaws had sat unnoticed for years or decades; the oldest dated back roughly 40 years.
  • The same announcement admits GLM-5.3 performed worse than western models on other, unnamed security and coding benchmarks.
  • Register columnist Simon Sharwood reads the result as evidence China closed the gap in offensive-security capability quickly after Anthropic's Mythos, eroding any US edge in the field.

Why it matters

A Chinese lab claiming parity, or better, with Anthropic and OpenAI specifically at finding software vulnerabilities is a notable capability marker, not just another leaderboard entry: bug-hunting AI has direct offensive and defensive security uses. Zhipu frames the result as a jump in reasoning across multi-stage exploitation chains rather than just spotting isolated flaws, which is the harder and more consequential skill.

Who it affects

Anthropic and OpenAI, whose models are the named benchmarks GLM-5.3 is measured against; security teams and open-source maintainers whose codebases may contain the kind of long-lived flaws Zhipu says GLM-5.3 can surface; and, per Sharwood's analysis, the broader US-China contest over AI-driven cyber capability.

How to use it

The article gives no price, availability details or access terms for GLM-5.3 beyond noting it launched last week and that Zhipu ran it against real-world code with unnamed Chinese company partners.

How solid is it

The headline claim rests entirely on Zhipu's own announcement and its own CyberGym benchmark run; no independent verification, exact score or margin is cited. The same announcement concedes GLM-5.3 trailed western models on other, unspecified security and coding benchmarks, which cuts against reading the CyberGym result as overall superiority.

Risks and caveats

An AI model that is unusually good at finding exploitable vulnerabilities is a dual-use capability: the same skill that helps defenders patch old bugs helps attackers find new ones first. The source does not say whether any of the 2,436 vulnerabilities found were disclosed, patched or assigned CVEs, nor which companies or projects were involved, so the practical handling of those findings is unknown.

“As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain.”

— Zhipu, in its GLM-5.3 announcement