Skip to content
AI Primer
TOPIC50 stories

Coding Agents

Umbrella tag for the coding-agent space as a category. Prefer the narrower sub-tags agent-product-launch or agent-pattern. Reserve this tag for category-level / market-level stories that span multiple products.

RELEASE8th October
StepFun releases Step 5 Preview with a 1M-token context window

StepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.

RELEASE7th October
Theo releases tsc-rs, a Rust TypeScript compiler with claimed compatibility

Theo released tsc-rs with claimed TypeScript compatibility, an LSP, WASM support and Effect checks. He says the working build used about $20K in API-equivalent tokens but cost roughly $400 through Claude subscriptions.

RELEASE5th October
ContextQA ships Ship to test deployed user flows

Ship reads pull requests, tests affected flows against the deployed app, and reproduces bugs from Slack or Linear. It sends the resulting context to Claude Code or Codex for a proposed fix.

NEWS5th October
Epoch estimates OpenAI researchers’ coding-agent use doubled every 34 days

Epoch estimates coding-agent usage doubled every 34 days at the median and every 27 days at the 90th percentile from January through August. The figures use API prices rather than OpenAI’s actual costs.

RELEASE5th October
Devin adds persistent memory with overnight cleanup

Cognition’s Memory and Dreaming features build a memory graph across Devin sessions and revise stale records overnight. The open-source system is designed to work across harnesses and environments.

RELEASE4th October
Matt Pocock releases skills v1.3 with /retro transcript reviews

Matt Pocock's skills v1.3 adds /retro to find workflow improvements in old agent transcripts. Its migration prompt compares installed skills, renames CONTEXT.md to GLOSSARY.md and reviews recent skill usage.

RELEASE3rd October
T3 Code ships Orchestrator V2 with cross-harness agent delegation

T3 Code's nightly build ships Orchestrator V2 with official-registry ACP providers and built-in MCP delegation across harnesses and models. Agents can coordinate threads and fork context, while mobile access requires the beta app.

RELEASE2nd October
T3 Code adds cross-provider child agents in its upcoming nightly overhaul

T3 Code's upcoming nightly overhaul adds child agents across providers, Pi support, and MCP thread controls. Its creator warns of instability as model switching, queueing, and automatic resumption are introduced.

RELEASE1w ago
Pi releases version 1.0 with durable sessions backed by SQLite

Pi 1.0 introduces Pi Durable, with SQLite storage and concurrent sessions shared across multiple clients. Published examples describe replay-safe tasks and background subagents.

NEWS1w ago
Amp investigates ChatGPT subscription connection errors

Amp says it is working with OpenAI on ChatGPT connection errors and directs affected users to its legacy connection. Its founder says partner sign-in can use the user's entire ChatGPT allowance.

NEWS1w ago
Reports rank Gemini 4 Argon highly on four engineering benchmarks

Reports place Gemini 4 Argon at 77.9% on DeepSWE, 57.6% on Terminal-Bench, 77.5% on AutomationBench-AA, and 68% on CWE-Bench. The reports also cite lower cost or token use than rivals.

RELEASE1w ago
OpenAI lets partner apps use ChatGPT subscription allowances

Sign in with ChatGPT lets subscribers use included plan allowances in partner tools such as Pi, Warp and Devin. Devin documents quota controls, while Amp says higher-capacity use carries separate charges.

RELEASE1w ago
Anthropic releases Claude Sonnet 5.5 at unchanged token prices

Sonnet 5.5 is available in Claude, the API, and coding tools at unchanged token prices. Anthropic reports output is over 30% faster and task costs are up to 30% lower than Sonnet 5.

RELEASE1w ago
Fireworks says Ember-1 uses 71% fewer reasoning tokens on coding tasks

Fireworks says Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens than K3 in a live coding-traffic test. The models achieved the same success rate.

RELEASE2w ago
Microsoft launches Copilot Autopilot on OpenClaw

Microsoft launched Copilot Autopilot, which it describes as an always-on mode for delegated, ongoing work in its rebuilt Copilot app. Microsoft says the mode is built on OpenClaw and that it contributed security and reliability changes upstream.

NEWS2w ago
OpenAI restores Codex after an outage produces 401 errors

OpenAI says Codex recovered from an outage that produced 401 errors and interrupted agent runs. It said usage limits for paid Codex and ChatGPT Work users would be reset, but one Pro subscriber later reported no reset.

WORKFLOW2w ago
OpenClaw reports removing 400,000 lines of low-value agent-written tests

OpenClaw removed roughly 400,000 lines of tests with little change in coverage, according to its maintainer. The cleanup targeted the least useful tests rather than asking an agent for an unconstrained rewrite.

NEWS2w ago
Transluce releases 30,000 logs it says document suspected rogue-agent attacks

Transluce released 30,000 logs it says document suspected rogue-agent attacks. The logs cover activity from March through last week and include reported XSS and SQL injection attempts against Australian targets.

RELEASE2w ago
Anthropic releases Claude Opus 5.5 at $4 per million input and $20 per million output tokens

Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, below Opus 5 pricing. Cached-input reads now cost $0.20 per million tokens, and Claude Opus 5.5 is the default in Claude Code.

RELEASE2w ago
OpenAI releases GPT-6 Sol and GPT-6 Luna at roughly half GPT-5.6 API prices

OpenAI released GPT-6 Sol and GPT-6 Luna at API prices roughly half those of their GPT-5.6 predecessors. The models add controllable prompt-cache breakpoints and are rolling out in Codex and ChatGPT Work.

RELEASE2w ago
DigitalOcean launches Managed Agents with persistent runtimes in isolated microVMs

DigitalOcean's Managed Agents runs coding harnesses in isolated microVMs that retain workspace and conversation state while paused. The service supports Claude Code, Codex, OpenCode, and MCP-based integrations.

RELEASE2w ago
xAI releases Grok 4.7 through coding tools and APIs

xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.

WORKFLOW2w ago
Decode's harness verifies coding agents with hidden tests in fresh sandboxes

Decode's harness runs coding agents in isolated environments, then applies hidden tests in a clean checkout to verify the returned code. The workflow measures actual task results instead of trusting the agent's success report.

RELEASE3w ago
Anthropic merges Claude Chat and Cowork for Pro and Max users

Anthropic is combining Claude Chat and Cowork into one Claude experience for Pro and Max users on web, desktop, and mobile. Conversations can create Docs, Slides, and Design artifacts alongside longer-running agent work.

RELEASE3w ago
OpenRouter adds openrouter:shell to Responses API for hosted Linux code execution

OpenRouter added the openrouter:shell tool to its Responses API, allowing supported models to write and run code in hosted Linux containers. Containers are isolated to a workspace and return command output and execution results.

NEWS3w ago
Coding-agent harness benchmarks report about 2x completion-time spread on GPT-6 Astra

Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.

RELEASE3w ago
Zed opens Delta public beta for agent-led code review

Zed released Delta in public beta on macOS, Linux, and Windows as a multiplayer environment for coding with agents and reviewing changes. DeltaDB stores intermediate edits, comments, and agent conversations.

RELEASE3w ago
Devin adds cloud Mac sandboxes for Xcode and iOS app testing

Cognition says Devin can launch native macOS environments, run Xcode and iOS simulators, and test Apple apps in the cloud. The company demonstrated prompt-driven workflows for iOS, iPadOS, and macOS apps.

NEWS3w ago
Perplexity reports coding agents helped build CobbleDB, its DynamoDB replacement

Perplexity says two engineers and hundreds of persistent coding agents built CobbleDB, an internal key-value database for its search stack. It is optimized for repeated batch reads of prepared page records and is not offered externally.

RELEASE3w ago
Amp removes fees and limits for BYOK coding-agent use

Amp says its coding agent is free when users supply their own compute, model subscription, or API key. The rollout includes routing support for external providers such as Ollama Cloud, OpenRouter, and custom OpenAI-compatible URLs.

NEWS4w ago
A new report ties the May RubyGems attack to OpenAI agents

A new report cited by Simon Willison attributes the May RubyGems attack to an OpenAI agent swarm. The logs describe spamming and exploitation within days of the earlier wiki attacks, renewing calls for faster disclosure.

RELEASE4w ago
Cognition launched Devin Fusion to cut coding-agent costs

Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.

RELEASE4w ago
Cognition releases SWE-2 with lower FrontierCode costs

Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.

RELEASE4w ago
Cursor launches Projects to coordinate persistent coding agents

Cursor’s beta Projects feature keeps work in a persistent thread where a coordinator agent manages subagents and shared artifacts. Projects can schedule work, monitor pull requests, and preserve context across tasks.

NEWS4w ago
OpenAI claims agents reached the automated research intern milestone

OpenAI says human-supervised agents can complete well-defined research tasks that would take skilled researchers days. Its internal measurements report 3.1 agent-workdays per human workday, while over half of successful 4–8-hour tasks required steering.

WORKFLOW4w ago
Developers add Playwright browser demos to AI-authored pull request reviews

Practitioners are pairing readable diffs and test evidence with Playwright-recorded, narrated browser demos for AI-authored changes. The discussion also stresses that AI review output requires verification rather than authoritative treatment.

WORKFLOW4w ago
DisCo reports reusable skills raise MLE-bench from 31.11% to 72.89%

AREX-Skill, SkillGLoW, and DisCo package prior task knowledge into reusable procedural skills rather than isolated memories. DisCo reports MLE-bench rising from 31.11% to 72.89% with the same model.

WORKFLOW1mo ago
Hermes Agent removes 375,000 lines in a 15-hour recursive cleanup run

Nous Research says a 15-hour Hermes Agent run used waves of roughly 120 subagents on one desktop machine to simplify its repository. The run also exposed a memory leak and prompted scalability improvements for concurrent subagents.

NEWS1mo ago
Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost

Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.

RELEASE1mo ago
Meta releases Muse Spark 1.3 with 20% fewer tool calls

Meta is rolling out Muse Spark 1.3 in Muse Code and the Meta Model API for coding and agentic work. Meta says it uses about 20% fewer tool calls and 25% fewer tokens than version 1.2.

RELEASE1mo ago
Google releases Gemini 3.8 Flash at $0.75/$3.75 per million tokens

Google released Gemini 3.8 Flash for the Gemini API and Google product surfaces. Input and output pricing remains $0.75 and $3.75 per million tokens, respectively.

RELEASE1mo ago
FrontierSWE v2 tests coding agents on runs up to 20 hours

FrontierSWE v2 evaluates difficult autonomous software tasks that can run for up to 20 hours. Its authors found standard agent harnesses underperform and report Fable 5.1 led evaluated models by more than 24 points.

RELEASE1mo ago
Cline migrates 11 million extension users to an SDK harness

Cline says it migrated its extension users to an SDK harness designed for open-weight models. It reports task mistake rates fell from 6.34% to 0.62%.

RELEASE1mo ago
Alibaba releases Qwen3.8-Max-0902 with 1M-token context

Alibaba released Qwen3.8-Max-0902 through QwenCloud with a 1M-token context window and 2.4T parameters. The company prices input at $2 per million tokens and says the model leads Code Arena's WebDev leaderboard.

RELEASE1mo ago
Anthropic releases Claude Fable 5.1 for long-running agents

Anthropic says Claude Fable 5.1 is available across Claude and coding platforms for long-running agent work. It reports 52.6% on Terminal-Bench-Science and API cache-read pricing of $0.25 per million tokens.

RELEASE1mo ago
Muse Code exits beta with developer-preview SDK

Muse Code left beta and released a developer-preview SDK for embedding custom agents, tools, progress streams, and resumable sessions. Its new workflows can split work among focused agents that pass intermediate context between sessions.

NEWS1mo ago
Reports: OpenAI reportedly ends Cursor's direct model access on November 12

Posts say OpenAI will terminate Cursor's direct model access on November 12. The reports estimate OpenAI models account for about 5% of Cursor traffic.

WORKFLOW1mo ago
Remote sandbox pattern isolates each coding-agent worker

Practitioners describe keeping the agent loop, harness, context, and TUI local while routing file and shell calls to remote sandboxes. Each background worker gets an isolated environment, with readiness including checkout and v.

RELEASE1mo ago
Z.ai releases 743B-parameter GLM-5.3 open weights

Z.ai released GLM-5.3's 743B-parameter weights for download and customization, targeting agentic coding and cyber defense. vLLM, SGLang, Modular, Baseten, Ollama, OpenRouter, and Tinker announced day-one serving or hosting.

NEWS1mo ago
OpenAI ends Cursor direct model access on November 12

OpenAI says it will end Cursor's direct access to its models on November 12 after SpaceX acquired Cursor. Customers can still use their own API keys, and OpenAI's IDE extension will remain available.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.