Kimi K3 and GLM-5.3 close in on Opus 5 and GPT-5.5

THE DECODER has published the fourth issue of its deep dive series Frontier Radar, arguing that Chinese open-weight AI models have closed most of the benchmark gap with Western frontier systems, leaving a measurable Western lead in only a few, increasingly narrow areas. Frontier Radar runs six times a year as a newsletter and an exclusive feature on THE DECODER's site for subscribers; earlier issues covered the state of AI agents, the measurable effects of AI on productivity, and the emerging token economy of AI. This fourth issue, the outlet says, is its most extensive yet, and sets out to cover how Chinese models caught up, where Western labs can still stand out, why the industry's moat has shifted, and why Europe is losing two races at once, though the last of those four threads falls outside the text available for this retelling.
The piece frames the shift against DeepSeek's R1, which caused a market shock about a year and a half earlier: a Chinese lab was suddenly competing with OpenAI's o1, the first commercial reasoning model, reportedly for far less money, and billions in market value evaporated within days. At the time, the picture was murkier than the headlines suggested. By DeepSeek's own report, R1 beat o1 on some individual tests, such as AIME 2024, but trailed clearly on others, including the factual-knowledge test SimpleQA, and later benchmarks showed the same pattern: Chinese models led in individual disciplines but not across the board. As recently as late June, when THE DECODER started work on this issue, Z.ai's GLM-5.2 still fit that older pattern.
Then came a new wave of Chinese open-weight releases: Moonshot's Kimi K3, Alibaba's Qwen3.8-Max, and Z.ai's GLM-5.3. THE DECODER says these models now sit near the top of almost every broad, demanding evaluation, handling long knowledge tasks, multi-step coding, and tool coordination far more reliably than their predecessors. Citing the Wall Street Journal, the outlet reports that Anthropic is fielding uncomfortable investor questions ahead of its planned IPO and is pointing to its remaining lead at the very top of the field in response; below that top tier, the piece says, the field belongs largely to cheaper, open Chinese models. Investors worry that raw model performance alone can no longer sustain a business: whatever a model can do exclusively today, a freely downloadable one may match within a few months.
THE DECODER lays out two accusations against the Chinese labs: that they used Western models as uncredited teachers through a technique called distillation, and that they tuned models to score well on benchmarks without matching broad capability, a practice it calls benchmaxxing. Its own position is that the question of guilt is almost beside the point. Whether or not the accusations hold, a model's lead cannot be defended once it is sold broadly, because any capability offered through an API can eventually be learned by a competitor. THE DECODER's stated thesis is that the industry's real moat is shifting away from any single model and toward the overall system that keeps producing the next one.
On the specifics: at launch, Kimi K3 placed third on Artificial Analysis's Intelligence Index with 57 points, just behind then-leaders GPT-5.5 and Opus 4.8, improving most on agentic tasks. On AutomationBench-AA, K3 briefly took first place until Anthropic answered with its newer Opus 5. On CEO-Bench, where an agent runs a fictional software company for 500 simulated days, K3 posted the best published single run at $22.15 million; its predecessor K2.7, like other Chinese models before it, had regularly failed these long-running tasks, and Qwen3.8-Max reaches a similarly high overall level. One caveat the piece raises: newer Chinese models sometimes burn far more tokens than Western ones on the same tasks, eating into part of their price advantage, so the cost-per-task math still applies.
Shortly after K3 launched, Opus 5 retook the top of the Intelligence Index with 61 points, a small gap over K3's 57, and THE DECODER identifies three areas where a Western lead is still measurable at all. The first is abstract, specialty testing. On ARC-AGI-1, a test of abstract pattern recognition using small puzzle grids, K3 and Fable 5, Anthropic's flagship line for coding and agent work, sit practically even at 94.5 and 98.5 percent, but the gap widens sharply on the harder ARC-AGI-2, to 60.4 versus 89.2 percent; THE DECODER cautions that ARC-AGI measures abstract reasoning far removed from everyday tasks, so this gap may not predict practical, economically relevant differences. The second is reliability. On the AA-AnalystAgent benchmark, launched August 12 to test agentic data analysis on real tables and documents, a task counts as solved only if a model gets it right in five out of five independent runs, a measure called pass^5; Opus 5 leads at 54 percent, GPT-5.5 at 50 percent, and K3, the best open model, at 39 percent. Yet K3 solves 73 percent of tasks at least once in five attempts, practically even with Opus 5's 74 percent, so the gap comes almost entirely from poor repeatability rather than raw capability. Reliability does not track the usual Intelligence Index rankings at all: GPT-5.6 Sol falls behind its own predecessor, GPT-5.5, on this measure. THE DECODER stresses that reliability matters commercially because an analyst agent only saves work when its answers hold up without review, and that running multiple passes or adding checkers can compensate but raises the cost of every accepted result.
The third and best-documented gap is cybersecurity. A joint assessment by the UK's AISI and the US CAISI found K3 far behind leading US models on offensive cyber tasks: 32 percent on the exploit-development test ExploitBench versus about 76 percent for the top US models on average; K3 failed all 41 tasks that required executing code on a target system, where the leading US models solved 20 on average; and in a simulated attack, K3 reached step 17 of 32 versus 28.5 on average for the US leaders. That gap is closing fast too: GLM-5.3, unveiled August 14, scores 54.4 percent on ExploitBench by Z.ai's own measurement, more than double predecessor GLM-5.2's score, cutting the distance to the US leaders roughly in half within a month; on CyberGym, a test of finding and validating vulnerabilities in source code, GLM-5.3 even edges past the leading US models. THE DECODER notes that cybersecurity is also the one area where providers withhold their strongest capabilities: Anthropic's restricted cyber model Mythos 5 hits 78 percent on ExploitBench but is only available under controlled conditions through Project Glasswing, while its public sibling Fable 5 effectively stays at the level of the older Opus 4.8, 40 percent, because upstream safeguards intercepted 407 of 410 test episodes. OpenAI follows a similar approach with Daybreak and other specialized cyber variants: GPT-5.6 Sol reaches 73.5 percent on ExploitBench in OpenAI's own evaluation but stays below Critical, the highest level on OpenAI's risk scale. With its upcoming Astra model, OpenAI says it cannot rule out for the first time that one of its own models crosses that threshold, which would trigger its Preparedness Framework more forcefully, up to a full development halt; OpenAI has already paused some internal Astra work. With GLM-5.3, THE DECODER says, a Chinese lab is adopting the same pattern for the first time: Z.ai is delaying the weights release by about two weeks for extra safety work and plans to limit the most sensitive cyber functions to verified users.
THE DECODER's synthesis: the abstract-testing gap is measurable but its economic value is unclear; the reliability gap matters most commercially but is the least settled, since rankings can flip even within one lab's own model line; and the cybersecurity gap covers capabilities that almost no customer can buy through regular channels anyway, and even that gap has narrowed sharply. Only a handful of the most dangerous capabilities can be walled off this way. The commercially valuable rest, coding, research, agent work, has to stay accessible through ordinary APIs or there is no business, and that is exactly where the lead shrinks within months.
The piece then turns to distillation itself. OpenAI and Anthropic accuse Chinese companies of using their models at scale to build competing ones. According to Anthropic, industrial campaigns it attributes to DeepSeek, Moonshot, and MiniMax jointly ran more than 16 million interactions through around 24,000 fraudulent accounts, with the campaign attributed to Moonshot alone targeting agent reasoning, tool use, coding, and reasoning traces across more than 3.4 million interactions. THE DECODER cites corroborating, if circumstantial, evidence: Together AI found a 0.72 correlation between the task-level success rates of K3 and Fable 5 on real software problems, and a style analysis by Typebulb placed K3 stylistically closest to Fable 5. Critics counter that the gap between Fable 5's release and K3's was too short to have influenced K3's training, but THE DECODER notes that Opus 4.8, a closely related Anthropic model, was available for much longer, and the same studies show clear overlaps there too. OpenAI, separately, describes a similar distillation pattern involving DeepSeek in its own memorandum to the US Congress. THE DECODER says plainly that none of the labs has published verifiable evidence for these accusations, and that its own inquiries to OpenAI and Anthropic turned up nothing new.
Drawing on the known accusations and outside research papers, THE DECODER reconstructs how such distillation could work mechanically. Training runs through pretraining on massive volumes of text, then midtraining on smaller, higher-quality data, then post-training through supervised fine-tuning and reinforcement learning, and distilled data can enter at nearly any of these stages: a leading Western model answers tens of thousands of selected tasks, complete with solution paths and tool calls, and the best of those answers can feed a student model's midtraining or fine-tuning, letting it pick up the teacher's solution patterns. For reinforcement learning, the teacher model does not even need to stay live: the collected data can be used to build a reward model instead. Anthropic's own report, per THE DECODER, shows this is not just a thought experiment: in the campaign it attributes to DeepSeek, Claude was made to process grading tasks at scale against predefined rubrics, effectively serving as a reward model for someone else's reinforcement learning.
Key facts
- Kimi K3 debuted third on Artificial Analysis's Intelligence Index with 57 points, just behind then-leaders GPT-5.5 and Opus 4.8, before Anthropic's Opus 5 retook the top spot with 61.
- On the ARC-AGI-1 puzzle test, K3 and Anthropic's Fable 5 are nearly tied (94.5 versus 98.5 percent), but the gap widens sharply on the harder ARC-AGI-2, to 60.4 versus 89.2 percent.
- On the pass^5 reliability benchmark AA-AnalystAgent, Opus 5 leads at 54 percent, GPT-5.5 at 50 percent, and K3 (the best open model) at 39 percent, yet K3 solves 73 percent of tasks at least once in five tries versus Opus 5's 74 percent, so the gap is mostly about repeatability, not raw skill.
- A joint assessment by the UK's AISI and the US CAISI found K3 well behind top US models on offensive cyber tasks (32 versus about 76 percent on ExploitBench), but GLM-5.3, unveiled August 14, has already more than doubled predecessor GLM-5.2's score to 54.4 percent by Z.ai's own measurement, cutting that gap roughly in half within a month.
- Anthropic says industrial distillation campaigns it attributes to DeepSeek, Moonshot, and MiniMax ran more than 16 million interactions through about 24,000 fraudulent accounts, including over 3.4 million tied specifically to Moonshot; neither side has published verifiable evidence either way.
Why it matters
The piece reframes a familiar story. A year and a half after DeepSeek R1's debut shocked markets by nearly matching OpenAI's o1 on some tests while trailing badly on others, the newest Chinese open-weight models, Kimi K3, GLM-5.3, and Qwen3.8-Max, no longer show that lopsided pattern: they now compete broadly, not just on isolated benchmarks. THE DECODER reports, citing the Wall Street Journal, that Anthropic is already fielding investor questions about this ahead of its IPO, because a model that only leads at the very top, while a much cheaper open alternative matches it lower down, is a harder business to defend. The outlet's own thesis carries the real weight here: whether Chinese labs got there through distillation of Western models or independent progress, a paid model's exclusive advantage now erodes within months once it is sold broadly. If the model itself cannot be the moat, the moat has to be the overall system that keeps producing the next one.
Who it affects
Investors and prospective buyers of Anthropic's stock are named directly, through the Wall Street Journal's report on IPO-related questioning. Enterprise buyers choosing between paid Western APIs and free or cheap Chinese open-weight models are the piece's implicit audience: it argues their case for paying a premium is narrowing to a few specific tasks. The Chinese labs it names, DeepSeek, Moonshot, and MiniMax, are directly affected by the distillation accusations raised by OpenAI and Anthropic, with OpenAI describing a similar pattern involving DeepSeek in a memorandum to the US Congress. The UK's AISI and the US CAISI supplied the piece's cybersecurity findings, and OpenAI itself says it has already paused some internal work on its upcoming Astra model over the same kind of risk threshold this piece discusses.
How to use it
Read as a buyer's guide, the piece's own numbers say the case for paying for a Western flagship model over a free Chinese open-weight one is now narrow and specific: abstract, puzzle-style reasoning, where the ARC-AGI-2 gap (60.4 versus 89.2 percent) is real and wide; work where getting an unreviewed answer right on the first try matters most, the pass^5 reliability gap measured by AA-AnalystAgent; and the handful of offensive-cyber capabilities Western labs keep behind controlled access anyway, such as Anthropic's Mythos 5 through Project Glasswing or OpenAI's Daybreak and other cyber variants. Outside those niches, the benchmarks THE DECODER cites show Chinese open models performing close enough on broad coding, knowledge, and tool-use tasks that a price premium is harder to justify. One practical offset the piece flags: newer Chinese models can burn far more tokens per task than Western ones, so the real cost-per-task comparison still has to be run rather than assumed from the sticker price.
How solid is it
Most of the comparative numbers rest on named third-party evaluators. Artificial Analysis is explicitly credited for the Intelligence Index; the UK's AISI and the US CAISI jointly ran the cybersecurity assessment behind K3's ExploitBench and simulated-attack scores; and ARC-AGI is an established, independently run test suite. The source does not say who operates or scores AutomationBench-AA, AA-AnalystAgent, or CEO-Bench. Some figures are self-reported instead: GLM-5.3's 54.4 percent ExploitBench score and its CyberGym result come from Z.ai's own measurement, and GPT-5.6 Sol's 73.5 percent comes from OpenAI's own evaluation, so those carry less independent verification than K3's AISI and CAISI figures. The distillation case is the weakest evidentially: Anthropic's interaction and account counts come from its own report, and the corroborating signals, Together AI's 0.72 correlation and Typebulb's style analysis, are correlational rather than direct proof. THE DECODER states plainly that none of the labs has published verifiable evidence either way and that its own inquiries to OpenAI and Anthropic turned up nothing new. The available text also runs out mid-explanation of how distillation could work mechanically, before any concluding assessment of the question.
Risks and caveats
THE DECODER flags several of its own limits. The ARC-AGI gap may measure abstract skill with no bearing on economically relevant tasks. Reliability rankings are volatile enough that a newer OpenAI model, GPT-5.6 Sol, trails its own predecessor, GPT-5.5. Critics quoted in the piece argue the release gap between Fable 5 and K3 was too short for distillation to have mattered, a point THE DECODER only partly rebuts by pointing to the longer-available Opus 4.8. On safety, OpenAI's own account is that it cannot rule out, for the first time, an internal model crossing its highest risk threshold, called Critical, with the upcoming Astra, which would trigger a fuller Preparedness Framework response, up to a full development halt; OpenAI has already paused some internal Astra work as a precaution. The distillation and benchmaxxing accusations against Chinese labs remain unproven by THE DECODER's own account and should be read as accusations, not established fact. Finally, this retelling reflects only the portion of the article available: its introduction also promises a section on why Europe is losing two AI races at once, which lies beyond the text retrieved here.
“The American lead hasn't disappeared. But it has retreated to a few, ever-narrower areas of the so-called frontier, the leading edge of what's technically possible.”
— THE DECODER, Frontier Radar #4