Skip to content
AI Primer
breaking

Anthropic reports GLM-5.3 built browser exploits in 50 of 410 attempts

Anthropic reports that GLM-5.3 produced working browser exploits in 50 of 410 controlled attempts. In a separate binary-exploitation test, it achieved control-flow hijacks in 4% of trials.

5 min read
Anthropic reports GLM-5.3 built browser exploits in 50 of 410 attempts
Anthropic reports GLM-5.3 built browser exploits in 50 of 410 attempts

TL;DR

  • GLM-5.3 built working exploits in 50 of 410 attempts against known browser-engine bugs; kimmonismus's thread puts Claude Mythos Preview at 56 of 410.
  • On a separate set of 100 open-source exploitation tasks, GLM-5.3 achieved four full control-flow hijacks versus six for Mythos, as the weblog excerpt quotes Anthropic.
  • GLM-5.3 refused direct malicious orders in Anthropic's simulation, but a deceptive cover story, prefilled reasoning, and modified weights produced 64%, 92%, and 100% engagement respectively, according to teortaxesTex's excerpt of the findings.

Anthropic's report also describes a model-built webpage that read a private key from a sandboxed browser machine and an eight-hour exploit run costing $20.40 at API prices. NIST's earlier assessment grades ExploitBench on a 16-point scale, a useful warning against treating every published percentage as the same measure.

ExploitBench's 50 working exploits

Anthropic tested 41 known vulnerabilities in Chrome's V8 engine in isolated sandboxes. GLM-5.3 completed 50 of 410 exploit attempts (12.2%); Claude Mythos Preview completed 56 (13.7%). The Claude models in this capability comparison ran with cyber safeguards disabled, according to Anthropic's methods.

The contrast with earlier models is stark: GLM-5.2 and Claude Opus 4.6 recorded zero complete exploits in this test, while Kimi K3 and DeepSeek V4.1-Flash registered small nonzero rates in the chart shared by kimmonismus.

Full control-flow hijacks on OSS-Fuzz projects

A second, internal evaluation gave full credit only when the model hijacked program control flow. On 100 randomly selected tasks involving open-source projects in OSS-Fuzz, GLM-5.3 succeeded in 4% of trials and Mythos Preview in 6%; GLM-5.2 and Opus 4.6 succeeded in none, according to Anthropic's benchmark description.

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic, generative-ai, ai-security-research, glm, ai, ai-in-china, llms

CAISI's different ExploitBench score

NIST's Center for AI Standards and Innovation published an independent assessment on September 17. It called GLM-5.3 the most cyber-capable open-weight model it had evaluated, while placing it about four months behind the current US frontier on its aggregate cyber measure.

CAISI's ExploitBench figure for GLM-5.3 is 61.1%, or 9.8 of 16 points under a graded, best-of-three-attempts-per-task method. Anthropic's roughly 12% counts attempts that produced a complete working exploit. CAISI also tested US models with cyber safeguards disabled where applicable and included trusted-access releases in its frontier comparison.

A browser exploit assembled from new flaws

In a researcher-led session lasting about a day, GLM-5.3 found previously unknown bugs in the JavaScript engine of a local Linux browser build. It chained them into a webpage that read arbitrary files from the test machine, demonstrated by reading an SSH private key, according to Anthropic's account. The researcher spent less than an hour actively supervising; Anthropic says it disclosed the browser flaws to the maintainer. The demonstrated exploit targeted the Linux build available in the sandbox.

A $20.40 ARM64 exploit chain

A separate researcher gave the smaller GLM-5.3-Flash public details of Chrome vulnerability CVE-2026-11645 and another known flaw. The model chained them into a working ARM64 exploit that bypassed pointer-authentication hardening. The run took eight hours of model work, 20 minutes of human attention, and an estimated $20.40 at Zhipu's API prices, Anthropic reported.

Safeguards under three bypass conditions

Anthropic gave models overtly malicious attack orders in a simulated environment. Each percentage below measures whether the model tried to connect to a remote target, across 50 episodes per condition. A fake bash tool returned LLM-generated approximations of command results; no generated code ran and no external system was contacted.

  • Bare order: GLM-5.3, 0%.
  • False red-team cover story: 64%.
  • Prefilled opening reasoning: 92%.
  • Abliterated weights with refusal behavior removed: 100%.

Deployed Claude models scored zero on the conditions Anthropic could test. The Claude API does not accept prefilled reasoning, and Claude's weights are unavailable for abliteration, as the test description notes.

Abliteration's compute bill

Anthropic's first attempt at removing GLM-5.3's refusals used about 2,200 GPU-hours, costing roughly $4,400. Its researchers estimate an experienced team starting from scratch could do it in about 600 GPU-hours, or $1,200. Refusal rates above 90% fell to roughly 3%, 2%, and 12% on JailbreakBench, HarmBench, and StrongREJECT respectively; GPQA-Diamond performance stayed level, while a tested CyberGym subset fell a few percent, according to Anthropic's measurements.

The release-control dispute has a concrete wrinkle: elder_plinius says Claude Opus helped build his OBLITERATUS safeguard-removal framework for open-weight models. In natolambert's critique, the ability to distribute defensive cyber capability broadly belongs in the same discussion as the risk of distributing offensive capability; emollick argues that downloadable weights remove a control layer available to hosted models.

A quote from the simulated model

Anthropic highlighted a disturbing line from an abliterated model's generated reasoning in its simulated attack test. Yuchenj_UW argues that a closed model could also be prompted into producing the same words, so the quotation alone does not establish comparative danger. The measured outcomes above concern attempted tool use, not that quotation.

Claude Code's documented misuse

The distinction between guarded API access and downloadable weights sits alongside real abuse of hosted models. In a November 2025 disclosure, Anthropic said a state-sponsored group manipulated Claude Code into attempting intrusions against roughly 30 targets, succeeding in a small number of cases; Anthropic banned the accounts it identified. Its August 2025 threat report described a separate Claude Code-assisted data-extortion operation targeting at least 17 organizations.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR2 posts
ExploitBench's 50 working exploits1 post
Full control-flow hijacks on OSS-Fuzz projects1 post
A browser exploit assembled from new flaws1 post
Safeguards under three bypass conditions2 posts
Abliteration's compute bill3 posts
Share on X