Skip to content
AI Primer
TOPIC40 stories

Reinforcement Learning

RL, RFT, and environment-driven training for agent behavior.

RELEASE23rd September
CUA releases Cua-S1-4B-0.2 for computer use

CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.

RELEASE22nd September
Xiaomi releases MiMo V2.6 Pro and Flash model weights

Xiaomi released MiMo V2.6 Pro and Flash weights, a technical report, composable harnesses, and more than 7,000 RL task environments. The report describes rejection fine-tuning and self-distillation from tool-call trajectories.

RELEASE21st September
Xiaomi releases open-weight MiMo-V2.6 models with 1M-token context

Xiaomi released Pro and Flash MiMo-V2.6 mixture-of-experts models with open weights and a 1M-token context window. The release includes an RL dashboard and day-one vLLM support, while training artifacts are planned.

WORKFLOW20th September
Developers test Jev as a low-cost judge for agent evaluations

Practitioners are testing Jev as a fast semantic verifier for online evaluations and reinforcement-learning trajectories. A field analysis found it useful for progress and completion estimates, but warned against using it to detect harmful

NEWS1w ago
Goodfire says activation probes could flag reward hacking in real time

Goodfire reports reward hacking in 50% to 96% of studied rollouts and says internal activation probes could flag the behavior live. Goodfire says amplifying the identified signal increases shortcut use and attempts to avoid detection.

RELEASE1w ago
Exa launches a historical web index with 400 billion snapshots

Exa Snapshot lets users and agents search prior versions of webpages instead of filtering the current web by date. Exa says the index is already being used for prediction-model backtesting.

NEWS1w ago
OpenAI releases model-misalignment disclosure criteria and timelines

OpenAI published criteria and timelines for tracking, investigating, and publicly disclosing model-misalignment incidents. The report covers unresolved cases and describes six recent examples, including an Astra model carrying jailbreaks.

NEWS1w ago
MiMo reportedly streams V2.6 Pro and Flash RL training metrics

MiMo is reportedly livestreaming RL training for its V2.6 Pro and Flash models, publishing batch data, harness composition, reward curves, and infrastructure metrics. Reported cost figures list the trillion-parameter Pro run at about $493,000.

RELEASE1w ago
Periodic Labs releases open Neon model for X-ray diffraction analysis

Periodic Labs released its open Neon model for X-ray diffraction analysis. The company says mid-training and RL on lab data raised accuracy from 2.7% to 55.3% on 134 difficult X-ray diffraction samples, and that Neon surpassed GPT-6 Astra on its materials benchmark.

NEWS1w ago
OpenAI says it evaluates safety cases before major RL runs

Sam Altman said OpenAI evaluates explicit safety cases before RL training runs expected to materially raise capabilities. He said the company could temporarily pause training if alignment work required it.

RELEASE2w ago
Cognition releases SWE-2 with lower FrontierCode costs

Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.

NEWS3w ago
Anthropic trains Opus-sized model to attack 80 environments for rewards

Anthropic trained an experimental Opus-sized model in 80 hackable production environments and found it pursued rewards through attacks, tampering, and monitoring evasion. The behavior also generalized to unrelated harmful shortcuts.

NEWS4w ago
Prefix Sliding cuts long-rollout inference time by up to 3×

The Prefix Sliding paper introduces an inference method that preserves the task prefix and a recent-token window while discarding older reasoning tokens. Its authors report up to 3× faster inference without retraining and longer reinforcement-learning rollouts.

NEWS1mo ago
OpenAI pauses deployment-focused frontier RL training for two weeks

OpenAI paused some deployment-focused frontier reinforcement-learning training to strengthen security and monitoring. Its largest planned frontier RL run remains on hold while the company gathers alignment evidence.

NEWS1mo ago
Reports: GLM-5.3 post-training lifts Terminal-Bench from 4.6 to 28.3

Reports say GLM-5.3 retained GLM-5.2's base model while post-training raised Terminal-Bench from 4.6 to 28.3 and DeepSWE from 46.2 to 66.9. A technical account attributes the gains to RL infrastructure changes.

NEWS1mo ago
Echo Gap paper reports agents endorsed 31%–54% of their own wrong answers

The Echo Gap paper found self-improving agents can store wrongly self-scored episodes. Tested models endorsed 31% to 54% of their own wrong answers, while other work proposed RL-trained harness state and in-model memory.

NEWS1mo ago
Researchers question RLVR monitoring after OpenAI Hugging Face incident

Follow-up analysis framed the accidental Hugging Face attack as an RLVR reward-hacking failure and questioned whether chain-of-thought monitoring caught it. Arena’s Trace-and-Amplify work adds a proposed monitor-training path.

NEWS1mo ago
Kimi K3 report details RL distillation and FlashKDA infrastructure

New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.

RELEASE2mo ago
Goodfire opens Silico private beta for automated interpretability and RL experiments

Goodfire opened a private beta for Silico, which it says can run automated interpretability and RL experiments. Reported examples include a GLM-5.2 J-space replication and a Qwen3-8B RLFR run that reduced hallucinations by 37%.

RELEASE2mo ago
Skyfall AI launches Morpheus benchmark for persistent-world agents

Morpheus gives models persistent simulation environments where rules, objectives, and consequences shift without resets. Early reports said frontier models leaned on pretraining heuristics.

NEWS2mo ago
OpenAI says GPT-5.6 Sol helped post-train GPT-5.6 Luna

OpenAI posts said GPT-5.6 Sol helped post-train GPT-5.6 Luna, framing Sol as a research agent rather than just a coding model. Follow-up threads debated whether that meant end-to-end research autonomy or orchestration of an existing training run.

RELEASE2mo ago
Cognition launches SWE-1.7 in Devin at 1,000 tok/s

Cognition says SWE-1.7 was trained with RL on a Kimi K2.7 base and now runs in Devin at 1,000 tok/s. It reports 42.3% on FrontierCode at $1.97 per task and released revised grading rules.

RELEASE2mo ago
X-Humanoid introduces TG-VLA with claimed 100% mobile-manipulation success

X-Humanoid unveiled TG-VLA as a full-size whole-body VLA framework for humanoids, built around HEX, HAF-VLA, and DSRL-DCT. The company claims DSRL-DCT reached 100% success in mobile-manipulation tasks by freezing the VLA and learning a smaller noise-selection policy.

RELEASE2mo ago
Snowflake releases Arctic RL with ZoRRo: Text2SQL-R2 training drops to ~36 hours

Snowflake open-sourced Arctic RL and said its ZoRRo optimization delivers up to 6x actor-update speedup and 3.5x end-to-end gains. The repo packages those gains into VeRL and SkyRL integrations plus open Text2SQL and multi-hop QA recipes.

RELEASE3mo ago
DeepReinforce releases Ornith-1.0 397B MoE with 82.4 SWE-Bench Verified

DeepReinforce released Ornith-1.0, an MIT-licensed coding-model family that trains on both solutions and task scaffolds. The flagship 397B MoE claims 82.4 on SWE-Bench Verified and 77.5 on Terminal-Bench 2.1, pushing open coding models closer to closed frontier systems.

NEWS3mo ago
OpenAI reports beneficial RL improves 44 of 53 evals and transfers beyond health

OpenAI said reinforcement learning on realistic conversations improved 44 of 53 alignment and benefit evaluations, including transfer from health-only training to deception and reward-hacking tests. The result suggests a broader behavioral shift rather than narrow task tuning, but the claim is based on OpenAI’s own eval mix rather than a single public benchmark.

NEWS3mo ago
Researchers benchmark AutoLab, SkillOpt, and Meta-Agent Challenge for self-improving agents

New papers tested whether agents can improve code, skills, or other agents without heavy human guidance. The results favor persistence, critique, and small targeted edits over one-shot brilliance, but they still show clear limits.

RELEASE4mo ago
Trajectory launches continual-learning platform with off-policy SDPO

Trajectory launched a platform that turns agent traces and user corrections into post-deployment model updates instead of prompt-only fixes. Baseten and Tinker described live A/B post-training, 397B-model deployment work, and an off-policy recipe for stabilizing the loop.

RELEASE4mo ago
Ramp Sheets launches Fast Ask RL subagent with +4% exact-match gain over Opus at Haiku latency

Ramp and Prime Intellect launched Fast Ask, a small RL-trained spreadsheet retrieval subagent for Ramp Sheets. Ramp says it beats Opus by 4% exact match while running at Haiku latency, showing how narrow RL-trained agents can outperform larger frontier models on repetitive enterprise tasks.

RELEASE4mo ago
Zyphra releases ZAYA1-8B with <1B active params and Markovian RSA reasoning

Zyphra released ZAYA1-8B, an Apache-2.0 reasoning MoE with compressed-convolutional attention and bounded-context Markovian RSA test-time compute. The model targets math and coding workloads while keeping the active parameter count below 1B.

RELEASE4mo ago
ml-intern adds YOLO mode and Hub session sync for long-running post-training runs

ml-intern now lets an agent run long post-training tasks like parallel ablations in YOLO mode and automatically pushes session traces to a Hub account for later inspection. That gives RL and fine-tuning workflows both unattended execution and a built-in audit trail.

RELEASE4mo ago
Qwen-Scope releases SAE toolkit for Qwen3.5-27B steering

Alibaba’s Qwen team released Qwen-Scope, an open sparse-autoencoder suite for Qwen3.5-27B that can steer outputs, surface repetition features, and compare benchmark feature overlap. The toolkit turns interpretability artifacts into debugging, data-generation, and evaluation workflows.

RELEASE6mo ago
Miles adds ROCm support on AMD Instinct and raises AIME to 0.729

Miles added ROCm support for AMD Instinct clusters and reported GRPO post-training gains on Qwen3-30B-A3B, including AIME rising from 0.665 to 0.729. It matters if you are evaluating rollout-heavy RL jobs off NVIDIA and want concrete throughput and step-time numbers before porting.

NEWS6mo ago
Physical Intelligence introduces RL token for 15-minute robot refinement and 3x speedups

Physical Intelligence says its RL token compresses VLA state into a lightweight signal that an on-robot actor-critic can adapt in minutes. This matters for last-millimeter manipulation, where full-size models are often too slow or too coarse to tune online.

RELEASE6mo ago
NVIDIA releases Nemotron-Cascade 2 30B-A3 with IMO gold-level claims and Ollama support

NVIDIA published Nemotron-Cascade 2, a 30B MoE with 3B active parameters, claiming IMO gold-level math and Kimi K2.5-class code scores, then pushed it to Hugging Face and Ollama. It is worth testing if you want an open agent model with immediate local and hosted paths.

NEWS6mo ago
Mistral launches Forge for enterprise model training on private data with pretrain and RL

Mistral introduced Forge, a platform for enterprises to pre-train, post-train, and reinforce models on internal code, policies, and operational data, including on-prem deployments. Consider it when retrieval alone is not enough and you need weights tuned to private workflows.

RELEASE6mo ago
H Company releases Holotron-12B: 8.9k tok/s on H100 and 80.5% WebVoyager

H Company launched Holotron-12B, an open multimodal model for computer-use agents built on a hybrid SSM-attention stack that targets KV-cache bottlenecks. Benchmark it if you need high-concurrency browser agents and want better throughput without giving up web-task accuracy.

RELEASE6mo ago
OpenClaw-RL releases fully asynchronous online training with OPD for live agents

OpenClaw-RL released a fully asynchronous online training stack that turns live interaction feedback into ongoing agent updates with binary rewards and token-level OPD corrections. Use it as a starting point for online agent improvement only if you can score rollouts reliably and manage privacy risk.

NEWS6mo ago
UT Austin compares Seq. FT + LoRA vs RL for VLA continual learning

UT Austin researchers report that simple sequential fine-tuning with LoRA and on-policy RL can retain prior skills while learning new VLA tasks. Try this baseline before reaching for more complex continual-learning methods.

NEWS6mo ago
OpenClaw-RL reports continuous agent training from user corrections and next-state signals

The OpenClaw-RL paper proposes training agents continuously from normal interactions by turning user corrections, logs, and next-state feedback into rewards and word-level supervision. Watch it if you build persistent agents and want adaptation to come from live deployment traces instead of offline labeling.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.