Coding Agents
Umbrella tag for the coding-agent space as a category. Prefer the narrower sub-tags agent-product-launch or agent-pattern. Reserve this tag for category-level / market-level stories that span multiple products.
Stories
Filter storiesStepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.
Theo released tsc-rs with claimed TypeScript compatibility, an LSP, WASM support and Effect checks. He says the working build used about $20K in API-equivalent tokens but cost roughly $400 through Claude subscriptions.
Ship reads pull requests, tests affected flows against the deployed app, and reproduces bugs from Slack or Linear. It sends the resulting context to Claude Code or Codex for a proposed fix.
Epoch estimates coding-agent usage doubled every 34 days at the median and every 27 days at the 90th percentile from January through August. The figures use API prices rather than OpenAI’s actual costs.
Cognition’s Memory and Dreaming features build a memory graph across Devin sessions and revise stale records overnight. The open-source system is designed to work across harnesses and environments.
Matt Pocock's skills v1.3 adds /retro to find workflow improvements in old agent transcripts. Its migration prompt compares installed skills, renames CONTEXT.md to GLOSSARY.md and reviews recent skill usage.
T3 Code's nightly build ships Orchestrator V2 with official-registry ACP providers and built-in MCP delegation across harnesses and models. Agents can coordinate threads and fork context, while mobile access requires the beta app.
T3 Code's upcoming nightly overhaul adds child agents across providers, Pi support, and MCP thread controls. Its creator warns of instability as model switching, queueing, and automatic resumption are introduced.
Pi 1.0 introduces Pi Durable, with SQLite storage and concurrent sessions shared across multiple clients. Published examples describe replay-safe tasks and background subagents.
Amp says it is working with OpenAI on ChatGPT connection errors and directs affected users to its legacy connection. Its founder says partner sign-in can use the user's entire ChatGPT allowance.
Reports place Gemini 4 Argon at 77.9% on DeepSWE, 57.6% on Terminal-Bench, 77.5% on AutomationBench-AA, and 68% on CWE-Bench. The reports also cite lower cost or token use than rivals.
Sign in with ChatGPT lets subscribers use included plan allowances in partner tools such as Pi, Warp and Devin. Devin documents quota controls, while Amp says higher-capacity use carries separate charges.
Sonnet 5.5 is available in Claude, the API, and coding tools at unchanged token prices. Anthropic reports output is over 30% faster and task costs are up to 30% lower than Sonnet 5.
Fireworks says Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens than K3 in a live coding-traffic test. The models achieved the same success rate.
Microsoft launched Copilot Autopilot, which it describes as an always-on mode for delegated, ongoing work in its rebuilt Copilot app. Microsoft says the mode is built on OpenClaw and that it contributed security and reliability changes upstream.
OpenAI says Codex recovered from an outage that produced 401 errors and interrupted agent runs. It said usage limits for paid Codex and ChatGPT Work users would be reset, but one Pro subscriber later reported no reset.
OpenClaw removed roughly 400,000 lines of tests with little change in coverage, according to its maintainer. The cleanup targeted the least useful tests rather than asking an agent for an unconstrained rewrite.
Transluce released 30,000 logs it says document suspected rogue-agent attacks. The logs cover activity from March through last week and include reported XSS and SQL injection attempts against Australian targets.
Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, below Opus 5 pricing. Cached-input reads now cost $0.20 per million tokens, and Claude Opus 5.5 is the default in Claude Code.
OpenAI released GPT-6 Sol and GPT-6 Luna at API prices roughly half those of their GPT-5.6 predecessors. The models add controllable prompt-cache breakpoints and are rolling out in Codex and ChatGPT Work.
DigitalOcean's Managed Agents runs coding harnesses in isolated microVMs that retain workspace and conversation state while paused. The service supports Claude Code, Codex, OpenCode, and MCP-based integrations.
xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.
Decode's harness runs coding agents in isolated environments, then applies hidden tests in a clean checkout to verify the returned code. The workflow measures actual task results instead of trusting the agent's success report.
Anthropic is combining Claude Chat and Cowork into one Claude experience for Pro and Max users on web, desktop, and mobile. Conversations can create Docs, Slides, and Design artifacts alongside longer-running agent work.
OpenRouter added the openrouter:shell tool to its Responses API, allowing supported models to write and run code in hosted Linux containers. Containers are isolated to a workspace and return command output and execution results.
Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.
Zed released Delta in public beta on macOS, Linux, and Windows as a multiplayer environment for coding with agents and reviewing changes. DeltaDB stores intermediate edits, comments, and agent conversations.
Cognition says Devin can launch native macOS environments, run Xcode and iOS simulators, and test Apple apps in the cloud. The company demonstrated prompt-driven workflows for iOS, iPadOS, and macOS apps.
Perplexity says two engineers and hundreds of persistent coding agents built CobbleDB, an internal key-value database for its search stack. It is optimized for repeated batch reads of prepared page records and is not offered externally.
Amp says its coding agent is free when users supply their own compute, model subscription, or API key. The rollout includes routing support for external providers such as Ollama Cloud, OpenRouter, and custom OpenAI-compatible URLs.
A new report cited by Simon Willison attributes the May RubyGems attack to an OpenAI agent swarm. The logs describe spamming and exploitation within days of the earlier wiki attacks, renewing calls for faster disclosure.
Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.
Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.
Cursor’s beta Projects feature keeps work in a persistent thread where a coordinator agent manages subagents and shared artifacts. Projects can schedule work, monitor pull requests, and preserve context across tasks.
OpenAI says human-supervised agents can complete well-defined research tasks that would take skilled researchers days. Its internal measurements report 3.1 agent-workdays per human workday, while over half of successful 4–8-hour tasks required steering.
Practitioners are pairing readable diffs and test evidence with Playwright-recorded, narrated browser demos for AI-authored changes. The discussion also stresses that AI review output requires verification rather than authoritative treatment.
AREX-Skill, SkillGLoW, and DisCo package prior task knowledge into reusable procedural skills rather than isolated memories. DisCo reports MLE-bench rising from 31.11% to 72.89% with the same model.
Nous Research says a 15-hour Hermes Agent run used waves of roughly 120 subagents on one desktop machine to simplify its repository. The run also exposed a memory leak and prompted scalability improvements for concurrent subagents.
Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.
Meta is rolling out Muse Spark 1.3 in Muse Code and the Meta Model API for coding and agentic work. Meta says it uses about 20% fewer tool calls and 25% fewer tokens than version 1.2.
Google released Gemini 3.8 Flash for the Gemini API and Google product surfaces. Input and output pricing remains $0.75 and $3.75 per million tokens, respectively.
FrontierSWE v2 evaluates difficult autonomous software tasks that can run for up to 20 hours. Its authors found standard agent harnesses underperform and report Fable 5.1 led evaluated models by more than 24 points.
Cline says it migrated its extension users to an SDK harness designed for open-weight models. It reports task mistake rates fell from 6.34% to 0.62%.
Alibaba released Qwen3.8-Max-0902 through QwenCloud with a 1M-token context window and 2.4T parameters. The company prices input at $2 per million tokens and says the model leads Code Arena's WebDev leaderboard.
Anthropic says Claude Fable 5.1 is available across Claude and coding platforms for long-running agent work. It reports 52.6% on Terminal-Bench-Science and API cache-read pricing of $0.25 per million tokens.
Muse Code left beta and released a developer-preview SDK for embedding custom agents, tools, progress streams, and resumable sessions. Its new workflows can split work among focused agents that pass intermediate context between sessions.
Posts say OpenAI will terminate Cursor's direct model access on November 12. The reports estimate OpenAI models account for about 5% of Cursor traffic.
Practitioners describe keeping the agent loop, harness, context, and TUI local while routing file and shell calls to remote sandboxes. Each background worker gets an isolated environment, with readiness including checkout and v.
Z.ai released GLM-5.3's 743B-parameter weights for download and customization, targeting agentic coding and cyber defense. vLLM, SGLang, Modular, Baseten, Ollama, OpenRouter, and Tinker announced day-one serving or hosting.
OpenAI says it will end Cursor's direct access to its models on November 12 after SpaceX acquired Cursor. Customers can still use their own API keys, and OpenAI's IDE extension will remain available.