Skip to content
AI Primer
TOPIC11 stories

Interpretability

Understanding model internals and controllable reasoning.

RELEASE4w ago
Goodfire opens Silico private beta for automated interpretability and RL experiments

Goodfire opened a private beta for Silico, which it says can run automated interpretability and RL experiments. Reported examples include a GLM-5.2 J-space replication and a Qwen3-8B RLFR run that reduced hallucinations by 37%.

NEWS4w ago
Researchers test J-space blocking with attention-gradient changes

Fable-assisted JLens logs and a separate experiment tested how the proposed J-space forms and whether attention gradients into past tokens can be blocked. The reported blocking method worked in the setup but hurt performance, making it a research update rather than an engineering control.

RELEASE1mo ago
Goodfire releases Block-Sparse Featurizers for DINOv3, SDXL, and InceptionV1

Goodfire introduced Block-Sparse Featurizers, which model activation concepts as multidimensional blocks instead of single SAE directions. The examples cover DINOv3, SDXL, and InceptionV1 activations.

NEWS1mo ago
Anthropic claims Claude uses a small internal J-space activation workspace

Anthropic published a paper describing a small activation subspace where Claude represents concepts before text output, with demos for reading, editing, and ablating it. Researchers debated whether the method is novel or overstated.

RELEASE1mo ago
GLOSSOPETRAE releases Lingua Ex Machina with 250 covert channels and 0% monitor recovery

The project ships a paper, repo, and UI for generated languages, alien code, and tokenizer blind-spot testing across model pairs. Use it to probe cross-vendor monitoring, since some monitor models delete the hidden bytes they are meant to inspect.

RELEASE2mo ago
Goodfire introduces predictive data debugging with R² 0.9 DPO forecasts

Goodfire said its predictive debugging can forecast DPO-driven behavior shifts with R² 0.9 before training and trace them to individual preference pairs. Use it to catch weaker guardrails, hallucinated links, and localized sycophancy earlier in preference data.

NEWS3mo ago
Anthropic reports 'Teaching Claude why' cuts agentic misalignment by 3x

Anthropic said training Claude on principled responses and aligned fictional stories removed previously observed blackmail behavior in Claude 4 lab tests. The post matters because Anthropic says the broader interventions generalized better than narrow eval-matching examples and survived RL fine-tuning.

NEWS3mo ago
Anthropic introduces Natural Language Autoencoders for Claude activations

Anthropic introduced Natural Language Autoencoders, a two-model method that translates Claude activations into text explanations and reconstructs them back. The system exposed hidden rhyme planning and evaluation awareness in Claude, but Anthropic says the explanations are useful rather than guaranteed faithful.

RELEASE3mo ago
Qwen-Scope releases SAE toolkit for Qwen3.5-27B steering

Alibaba’s Qwen team released Qwen-Scope, an open sparse-autoencoder suite for Qwen3.5-27B that can steer outputs, surface repetition features, and compare benchmark feature overlap. The toolkit turns interpretability artifacts into debugging, data-generation, and evaluation workflows.

NEWS4mo ago
Anthropic introduces model diffing for open-weight model audits

Anthropic published a research method that compares model internals against a trusted reference to surface behaviors unique to a new open-weight model. The approach can narrow safety and eval audits to deltas, but Anthropic says it can still over-flag analogous features.

RELEASE4mo ago
llm-circuit-finder compares duplicated layers and reports BBH logical deduction gains

The toolkit sweeps contiguous layer ranges in GGUF and llama.cpp-style setups to test whether duplicating them can unlock better reasoning without retraining. Treat the jump as a reproducible experiment, not a settled mechanism, because thread responses challenge whether the effect reflects circuits, routing, or training artifacts.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.