Coding Agents
Stories about coding-agent products and patterns, including Claude Code, Codex, Cursor, and similar.
Stories
Filter storiesAnthropic has moved Claude Code cloud sessions out of research preview, allowing tasks to keep running on Anthropic-hosted infrastructure after a user closes their computer. Sessions use a Pro or Max plan, and Pro subscribers receive trial credit for the feature.
Posts report that GPT-6 Sol is live in Codex, Work, and the API alongside the lighter GPT-6 Luna. Reported API prices are roughly half those of the prior Sol models.
Anthropic says Claude Opus 5.5 is now the default across Claude Code, Claude, and Cowork for paid plans. The company says it is faster and cheaper than Opus 5, with rate limits extending 25% further.
xAI says Grok 4.7 improves on 4.6 without changing price or speed. It is available in Grok Build, Cursor, and the API, with higher scores on several agent benchmarks.
Several Codex subscribers say their paid allowances are falling rapidly during ordinary coding work. One user reported consuming 20% of a $200 plan in a day, while another reported 40% spent on a separate project.
Agent Picnic puts more than 100 AI agents—including Muse, Grokbot, ChatGPT, Claude, and Codex—in a shared chat with a coordinating agent. Its creator says the XMTP-based system can assign tasks, fact-check responses, and compare participants' work.
Atria Dawn Preview is presented as a research agent that validates results against tests, metrics, rendered files, file states, and citations. Its reported API has a 256K text context window and is text-only.
The workflow gives GPT-6 Astra authenticated MCP access to Hyper3D Rodin for planning and generating 3D assets. A related setup sends an image through Astra and Codex to create an editable 3D space rather than a static model.
Developer launch reports say the public-beta API supports long-running cloud agents, context compaction, parallel subagents, MCP, and sandboxes. It can run on hosted or self-managed infrastructure.
Claude Code’s `plugin eval` command runs test cases with and without a plugin, scores both runs, and produces terminal and HTML comparisons. Anthropic says developers can recheck skills after model releases, though evals consume tokens.
Amir Mushich shared a car-site demo made without 3D models using GPT-6 Astra, LTX video transitions, and GPT Image 2.5 for colors and wheels. Mushich also published the prompt and repository, making the workflow reproducible.
Higgsfield demoed GPT-6 Astra creating interactive 3D experiences including a wardrobe, room planner, Tokyo rail tracker, and hand-tracked story theatre. A Spline robot demo adds collision detection and pose correction to its walking solver.
Creators have shared browser games, Godot projects, simulations, and a released tower-defense game made or finished with GPT-6 Astra. Reported workflows combine Astra with Blender, Three.js, Fable, Godot, and imported 3D assets.
Amir Mushich published a workflow for a web interface where clicks move between generated video states. It uses planned keyframes, matched forward and reverse clips, and a direct LTX API connection.
Developers report publishing playable web and 3D games built with GPT-6 Astra. Reported projects include a four-level Pitfall clone and Blender/Godot builds; one tutorial covers four games made in about two hours.
Google released Gemini 3.8 Flash across its developer and consumer products. Google says it matches Gemini 3.7 Flash pricing and scored 73.7% on DeepSWE 1.1.
Fable 5.1 is now available in Claude Code and the Claude Platform at Fable 5 pricing. Anthropic also cut API cache-read pricing from $1 to $0.25 per million tokens and reset five-hour and weekly user limits.
Figma Community now accepts plugins and shaders built with Config. The update also adds interactive canvas shaders, shader-code access, and MCP read/write support.
OpenClaw says users have moved to its web interface and native apps after a major rewrite. The project acknowledges one report that its migration tool missed legacy exec rules.
AIandDesign says its AI-built system documents 22 component families with synchronized JSX, tokens, accessibility, and testing guidance. A separate library offers 2,000 model-readable DESIGN.md files.
A Base44 user reports that a one-paragraph prompt produced a live CRM with authentication, persistent data, and a custom domain. They later added email follow-up reminders through another prompt.
A Tembo user demonstrated real-time frontend iteration in cloud Linux VMs with databases in the same environment. Apps can be tunneled from a VM and shared by link, while nested virtualization is being added.
JFrog released the free Boost CLI, which compresses shell output before it reaches coding agents including Claude Code and Codex. Its Terminal-Bench 2.0 report claims a 13.5% cost reduction across 89 tasks.
The MIT-licensed service says 58 builders with 56 sites are trading contextual links through coding agents. It gives an agent a partner brief, then checks whether the agreed link was published after repository edits.
A free tool connects coding agents to reciprocal-link matches, then lets them edit repositories, deploy changes, and verify placement. Its creator reports 58 builders and 56 sites participating across 19 categories.
Meng To released ThreeUI, an open-source library of procedural 3D landing-page components, icons, motion designs, and templates. Creators can copy its source or prompts into coding agents, while some extras remain paid.
A roundup of user examples says Grok Bot can build games, manage repositories through recruited agents, SSH into a Mac, buy domains, and deploy projects. The reports describe computer-use workflows extending from code generation to deployment.
Remotion Skills 2.0 simplifies its APIs, generates interactively editable code in Studio, routes best-practice prompts to relevant skills, and removes embedded design defaults. Studio output remains editable, while projects no longer receive built-in design defaults.
Cursor’s Origin hosts repositories, reviews pull requests, runs agents, and deploys through Vercel. GitHub repositories can sync while GitHub remains the source of truth.
Users report GPT 5.6 Sol's selected 1M-token context window initially reset to roughly 258K after a message. A later user confirmation said the 1M setting was restored.
Developers described brittle AI app workflows, including one CSS-to-Tailwind refactor that had to be resumed five times. The complaints focused on debugging, local setup, and project inventory.
Alibaba released Qwen3.8-Max, described in a launch thread as a 2.4T sparse MoE with 95B active parameters, 1M context, native vision/text, and agent benchmarks. API pricing is listed at $2/$6 per 1M tokens, with open weights planned for Hugging Face.
Boris Cherny said Anthropic has landed large P99 RSS improvements for Claude Code while still collecting CPU and memory reports across CLI and Desktop. The issue remains an active performance investigation requiring machine, OS, version, and task details rather than a fully verified fix.
Anthropic says Claude Code Pro, Max, and Team users will default to Auto Mode on Aug. 14. Its tool-call classifier reportedly caught 89% of dangerous commands, versus 14% for manual approval, after prompt-injection testing.
Graft scans a repo into linked markdown and plugs into Claude Code hooks so the agent can reuse codebase context. Its 162-run benchmark claims 46% fewer tool calls, up to 4x fewer tokens, and 60% less time.
Developers proposed agent-specific browser auth for Codex, Devin, Claude, and Cursor so agents can use selected passkeys, TOTPs, and passwords without exposing full vaults. Related posts warned that Claude Connectors may broaden tool access.
Developers shared workflows that turn idle AI accounts into scheduled assistants for commits, research, repo cleanup, and notes. Shared designs add spend caps, Apple Notes routing, Obsidian output, and coworker controls.
Meng To released code and prompts for a scrolling bookshelf demo refined through many Codex prompts. Karpathy separately used Claude Opus 5 to turn the opening of LOTR into a 5,500-line Three.js render.
Practitioners said coding agents can make impressive game demos quickly, but playable games still require weeks or months of design, testing, and cleanup. Other posts pointed to limited visual reasoning, 3D topology fixes, and Fable or Opus model choices as current bottlenecks or workarounds.
Practitioners are putting review steps around AI app builders before merge or publish. Examples include PostHog-to-Cursor Cloud PRs, Bolt's six-category security scan, and a Claude Code prompt that inspects repos first.
Pieter Levels summarized Fireship’s claim that cheap coding agents make micro-SaaS execution easier to copy. Replies argued distribution, complexity, and non-code moats still matter.
LLMJunky demoed Slopet League, a Rocket League-style prototype in Godot, saying Opus 5 made the map, physics, effects, and animations. He says GPT 5.6 Sol made car meshes and the build took about 1.5 days.
Posts linked to Kimi K3’s weight release, with creators testing it in Claude Code, Cursor, and local hardware discussions. Aakash Gupta said Moonshot revenue rose at least 6x after launch.
Peter Yang shared Jason Liu’s workflow for mining past sessions into skills, adding a /write-like-me skill, pinning durable threads, importing browser sign-ins, and bundling related skills as plugins. One example edited a launch video from Slack feedback.
Creators shared playable game prototypes built with coding agents, including Codex rebuilding Octane in Blender, a Red Alert clone, a Three.js arena shooter, and a Claude Code tank game. Other demos used single-file browser-game scaffolds.
New Opus 5, Codex, and Fable demos showed agents creating 3D worlds, Blender assets, Godot ports, and one-shot games. Pushback focused on whether the clips become games worth playing.
Shared workflows used Codex, Claude Code, and Fable for release testing, inbox triage, product-plan review, and adversarial code review. OpenClaw’s QA prompt split testing across 12 subagents.
Shared resources included design-skills repos, a UI visual dictionary, CRM table tokens, and Claude Opus 5 website prompts. The reusable context can speed front-end generation but does not replace product-specific design choices.
Anthropic launched Claude Opus 5 for Claude Code and the API with high-effort defaults, fast mode, migration tooling, mid-conversation tool changes, and routing fallback. The model scored 43.3% on Frontier-Bench v0.1.
Creators posted demos where Opus 5 built browser 3D games, Codex Sites produced a Three.js game, and Grok 4.5 helped redesign a Steam card game. Some workflows now pair video concepts with Claude Code, but the evidence is demo-based.