Coding Agents
Stories about coding-agent products and patterns, including Claude Code, Codex, Cursor, and similar.
Stories
Filter storiesAlibaba released Qwen3.8-Max, described in a launch thread as a 2.4T sparse MoE with 95B active parameters, 1M context, native vision/text, and agent benchmarks. API pricing is listed at $2/$6 per 1M tokens, with open weights planned for Hugging Face.
Boris Cherny said Anthropic has landed large P99 RSS improvements for Claude Code while still collecting CPU and memory reports across CLI and Desktop. The issue remains an active performance investigation requiring machine, OS, version, and task details rather than a fully verified fix.
Anthropic says Claude Code Pro, Max, and Team users will default to Auto Mode on Aug. 14. Its tool-call classifier reportedly caught 89% of dangerous commands, versus 14% for manual approval, after prompt-injection testing.
Developers proposed agent-specific browser auth for Codex, Devin, Claude, and Cursor so agents can use selected passkeys, TOTPs, and passwords without exposing full vaults. Related posts warned that Claude Connectors may broaden tool access.
Graft scans a repo into linked markdown and plugs into Claude Code hooks so the agent can reuse codebase context. Its 162-run benchmark claims 46% fewer tool calls, up to 4x fewer tokens, and 60% less time.
Developers shared workflows that turn idle AI accounts into scheduled assistants for commits, research, repo cleanup, and notes. Shared designs add spend caps, Apple Notes routing, Obsidian output, and coworker controls.
Practitioners said coding agents can make impressive game demos quickly, but playable games still require weeks or months of design, testing, and cleanup. Other posts pointed to limited visual reasoning, 3D topology fixes, and Fable or Opus model choices as current bottlenecks or workarounds.
Meng To released code and prompts for a scrolling bookshelf demo refined through many Codex prompts. Karpathy separately used Claude Opus 5 to turn the opening of LOTR into a 5,500-line Three.js render.
Practitioners are putting review steps around AI app builders before merge or publish. Examples include PostHog-to-Cursor Cloud PRs, Bolt's six-category security scan, and a Claude Code prompt that inspects repos first.
Pieter Levels summarized Fireship’s claim that cheap coding agents make micro-SaaS execution easier to copy. Replies argued distribution, complexity, and non-code moats still matter.
LLMJunky demoed Slopet League, a Rocket League-style prototype in Godot, saying Opus 5 made the map, physics, effects, and animations. He says GPT 5.6 Sol made car meshes and the build took about 1.5 days.
Posts linked to Kimi K3’s weight release, with creators testing it in Claude Code, Cursor, and local hardware discussions. Aakash Gupta said Moonshot revenue rose at least 6x after launch.
Peter Yang shared Jason Liu’s workflow for mining past sessions into skills, adding a /write-like-me skill, pinning durable threads, importing browser sign-ins, and bundling related skills as plugins. One example edited a launch video from Slack feedback.
Creators shared playable game prototypes built with coding agents, including Codex rebuilding Octane in Blender, a Red Alert clone, a Three.js arena shooter, and a Claude Code tank game. Other demos used single-file browser-game scaffolds.
New Opus 5, Codex, and Fable demos showed agents creating 3D worlds, Blender assets, Godot ports, and one-shot games. Pushback focused on whether the clips become games worth playing.
Shared workflows used Codex, Claude Code, and Fable for release testing, inbox triage, product-plan review, and adversarial code review. OpenClaw’s QA prompt split testing across 12 subagents.
Shared resources included design-skills repos, a UI visual dictionary, CRM table tokens, and Claude Opus 5 website prompts. The reusable context can speed front-end generation but does not replace product-specific design choices.
Anthropic launched Claude Opus 5 for Claude Code and the API with high-effort defaults, fast mode, migration tooling, mid-conversation tool changes, and routing fallback. The model scored 43.3% on Frontier-Bench v0.1.
Creators posted demos where Opus 5 built browser 3D games, Codex Sites produced a Three.js game, and Grok 4.5 helped redesign a Steam card game. Some workflows now pair video concepts with Claude Code, but the evidence is demo-based.
Anthropic’s developer update adds per-agent effort settings, seeded session creation, webhooks, sub-agent event streaming, and up to 500 skills per session. The release gives teams finer controls for managed agent runs.
Levelsio pointed to customers and founders replacing paid apps with private AI-built versions. The debate centers on whether AI raises quality baselines while weakening pricing power for simple software.
Google says Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber target faster, cheaper agent workloads. Early tests cite lower token use and stronger coding results.
Peter Yang shared Thariq’s explanation that Anthropic cut Claude Code’s system prompt by 80% as newer models need fewer examples and more room to reason. The same discussion covered /loop, /goal, workflows, and long-running agents.
ClaudeDevs said Pro, Max, Team, and seat-based Enterprise users keep 50% higher Claude Code weekly limits through August 19. Users had warned the promo’s end would sharply cut effective usage.
LLMJunky described a ChatGPT Work/Codex setup where agents in parallel threads exchange session IDs to work on the same codebase, wait, and request updates. Users also asked for better natural-language search across session history.
Levelsio shared a repeatable OpenCode setup using a paid Kimi Code API key and BUILD mode after OpenRouter limits and Claude Code safety blocks. Other users reported Kimi Code quotas interrupting small game builds.
Steipete said he moved the Clawsweeper GitHub review bot to 5.6 Terra high and saw about 40% faster reviews at lower cost with little quality change. A follow-up framed the result as a warning that general model benchmarks may not predict issue-review performance.
Creators posted open-weights benchmarks and tests comparing Kimi K3 with GPT-5.6 Sol and Fable 5. Demos covered UI animation, games, Seedance renders, kernel code, and reported OpenRouter rate limits.
Min Choi collected early GPT-5.6 Sol builds spanning Blender scenes, After Effects automation, Three.js characters, UI cloning, SQL DOOM, and Claude Code tests. Meng To also published a video-to-HTML workflow.
Min Choi's workflow assigned Grok 4.5 to research and debugging, Fable 5 to planning and frontend, and GPT-5.6 Sol to complex coding. Shann Holmberg added tactics for effort levels and usage limits.
Steipete said his Mac Studio hit its session limit, so he spreads coding-agent work over about five machines through Jump Desktop. He relies on autoreview and tests to catch mistakes.
Hasantoxr documents AgenticSeek with Ollama, SearXNG, and Docker, including install steps, model config, and a locked WORK_DIR. The stack keeps models, chats, and files on the user's machine.
Meng To says Sol Ultra generated an HTML/React canvas with layers, inspector panels, media items, and real-time collaboration. The test used planning and subagents, but it burned through limits.
Min Choi's roundup cites Grok 4.5 building a playable FPS, a UE5 cyberpunk street, a home planner, and tool-using agent setups. Viktor Oddy also posted a Cursor website tutorial using the model.
Users reported the former Codex Superapp now appears as ChatGPT Work, with Codex inside the ChatGPT desktop flow. Early posts flagged broken sites support and confusing sidebar states.
Levelsio used Claude Code on a VPS to SSH into a cloud Mac Mini, build an iOS app headlessly, and stream serve-sim into a browser. Follow-up posts documented the tunnel and serve-sim setup, making the workflow reproducible for remote iOS testing.
xAI announced Grok 4.5 as a coding-and-agent model available in the SpaceXAI console, Grok Build, and Cursor. Posts cite $2 per million input tokens, $6 per million output tokens, and delayed EU availability.
Kitze described using a $60 Hetzner server, four $200 Codex accounts behind codex-lb, self-hosted Paperclip workspaces, and a GPT-5.5 manager to run about 30 isolated tasks. The setup routes parallel coding work across isolated workspaces and paid Codex accounts.
Meng To showed Taste with Codex or Claude Code rebuilding landing pages from screenshots and videos. Voila also launched a tool that inspects UI and turns edits into code, copy, prompts, tasks, or agent handoffs.
Posts report Fable 5 remains included in Claude only through July 7 before moving to $10/$50 per million-token credits. Watch usage closely: tests show roughly 2x Opus consumption in subagent and one-shot app workflows.
OpenClaw's maintainers say the community-built iOS and Android apps are live, with secure pairing and push notifications. The rollout also exposed staffing, funding and plugin-support limits, so the project is still looking for contributors.
ClaudeDevs said it raised 5-hour and weekly usage caps for the weekend across every plan. The change lands as users report token burn, overage charges, and heavier Claude Code sessions driven by long agent runs and goal-checked workflows.
Users documented Codex handling self-service signups, repo-maintenance loops, and folder overwrite failures on June 14. Watch the wrapper update closely, since it also added rate-limit reset banking and browser dev mode around the same workflow.
Anthropic said a US government directive forced it to disable Fable 5 and Mythos 5 across Claude products and APIs. The change also pushed Build Day and downstream tooling to Opus 4.8, breaking active Fable sessions and triggering fallbacks in tools like Linear Agent.
New creator demos pushed Claude Fable 5 into CAD, landing pages, and web game ports, including an Autodesk Fusion RC car built in three prompts. Watch for longer runs to trip safeguards and fall back to Opus 4.8.
Grok Build opened a beta plugin marketplace and added official MongoDB, Vercel, and Sentry integrations alongside Chrome DevTools and Cloudflare support. The beta turns deployment, database, debugging, and observability tools into installable terminal actions.
Creators pushed Claude Fable 5 into browser platformers, a GTA 2 clone, a SNES port, CAD models, and a webcam fruit-slicing game. The demos show playable prototypes can now come from prompts plus a few follow-up fixes, so creators can move faster from idea to test build.
Anthropic opened Claude Fable 5 across Claude Code, Desktop, Cowork, and API, with always-on reasoning and Opus 4.8 fallback on some flagged requests. Early demand triggered model-picker friction and quota pressure, so Anthropic reset the 5-hour and weekly limits the same day.
Community demos showed Claude Fable 5 generating playable games and simulations from short prompts, image refs, and goal-based instructions, from Pokémon and F-Zero to city sims and FPS clones. The demos make the model’s creative ceiling clearer, but builders still needed follow-up prompts for speed, style, or polish.
Anthropic staff said Claude Code usage has shifted toward auto mode, routines, and phone-based coding one year after GA, and they pointed users to /usage for token breakdowns. The thread matters because it shows Anthropic’s intended daily-driver workflow as community comparisons with Codex intensify.