Silico
Automated interpretability and RL experimentation platform
Interpretability-first AI research platform that automates experiments, including causal interpretability replications and reinforcement-learning workflows, to inspect, debug, and shape model behavior.

Recent stories
Goodfire describes probes that inspect model internals for prohibited intent, reward hacking, and risky tool calls. Goodfire reports that a probe for a GLM 5.3 coding agent caught 99% of prohibited actions in its test.
Goodfire opened a private beta for Silico, which it says can run automated interpretability and RL experiments. Reported examples include a GLM-5.2 J-space replication and a Qwen3-8B RLFR run that reduced hallucinations by 37%.
Goodfire said its predictive debugging can forecast DPO-driven behavior shifts with R² 0.9 before training and trace them to individual preference pairs. Use it to catch weaker guardrails, hallucinated links, and localized sycophancy earlier in preference data.