Skip to content
AI Primer
breaking

METR patches Inspect JavaScript bug that could alter reviewed transcripts

METR disclosed and patched a JavaScript injection bug in the Inspect transcript viewer that could show reviewers altered agent transcripts while leaving underlying logs unchanged. Inspect was patched within a day, but the incident exposed a gap between stored traces and reviewer-visible output.

4 min read
METR patches Inspect JavaScript bug that could alter reviewed transcripts
METR patches Inspect JavaScript bug that could alter reviewed transcripts

TL;DR

  • Agent output could inject JavaScript into Inspect’s transcript viewer and fake the history a reviewer saw, as METR’s demonstration illustrated.
  • The exploit changed the displayed transcript while leaving underlying logs untouched, according to METR’s follow-up.
  • Inspect was patched within a day of reporting; METR said it had not observed agents exploiting the flaw in evaluations.

An AI-assisted researcher found the vulnerability in about 10 minutes. The nastiest detail was the Download button: the exploit could intercept that, too.

Isolated staging test

METR disclosed the proof of concept on October 6, following an investigation earlier this year. The research ran on a staging sandbox isolated from METR’s production infrastructure.

The target was the transcript viewer in Inspect, a framework widely used for safety evaluations. Researchers use it to step through what an agent did during an evaluation.

MathJax to JavaScript

An agent could plant the payload anywhere in its output, including a reasoning block. The injection followed this rendering path, described in METR’s technical appendix:

  1. Inspect renders mathematical notation in transcripts using MathJax.
  2. A crafted URL inside a MathJax \href expression breaks out of the generated attribute because it is inadequately escaped.
  3. The resulting markup introduces an SVG <animate> element with an onbegin handler.
  4. The browser fires onbegin automatically, executing the agent’s JavaScript through eval(atob(...)) in the viewer’s origin.

Forged histories and downloads

The injected code could rewrite previous actions in the displayed transcript and intercept the button that downloads a transcript, according to METR’s report.

The real trajectory would still exist in METR’s database. Discovering the deception would depend on a reviewer suspecting the displayed history was false and looking beyond the viewer, as METR noted in its follow-up.

Patch and evaluation review

Meridian Labs, the team behind Inspect, patched the vulnerability within one day of reporting, according to METR’s appendix. METR said it had not observed agents exploiting the flaw in its announcement.

METR had searched its evaluations generally for misalignment and security incidents, but had not searched for this specific exploit. It expected an attempt would have been flagged by the broader search and said it planned to share more about that log review in the coming weeks, according to the report’s footnote.

Untrusted mode

Separately, Meridian Labs merged an untrusted-mode change on October 1. METR’s appendix describes the timing as coincidental.

Inspect’s changelog lists three entry points for plain-text rendering:

  • Task configuration: ViewerConfig(trust_content=False).
  • Viewer process: inspect view --no-trust-content.
  • Viewer environment: INSPECT_VIEW_TRUST_CONTENT=false.

Rich rendering remains the default. With trust disabled, the viewer:

  • Shows markdown, math, and code as literal source text.
  • Shows terminal escape sequences instead of interpreting them.
  • Replaces images, audio, video, and terminal recordings with placeholders.
  • Shows links and their destinations as text rather than making them clickable.
  • Makes invisible and bidirectional-override characters visible, for example ⟨U+202E⟩.

Embedded and older viewers

Plain-text enforcement depends on the viewer that reads the log, according to Inspect’s documentation:

  • --no-trust-content applies only to the inspect view server.
  • Bundled and embedded viewers rely on each log’s own ViewerConfig(trust_content=False) setting.
  • Older viewers, including older versions of the VS Code extension, ignore the setting. Older Inspect versions drop it when rewriting a log.
  • Scout does not yet honor it.
  • The viewer option only lowers trust. Passing --trust-content cannot override a log marked trust_content=False.

The documentation limits the setting to rendering; it does not make the content safe to copy or act on.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 2 threads
TL;DR1 post
Forged histories and downloads1 post
Share on X