Computer Use
Agents that click, type, browse, and operate software directly.
Stories
Filter storiesCUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.
Meta said Muse for Mac can queue a job and continue operating a user's laptop without the user at the keyboard. The Connect rollout also gives Muse agents email addresses and adds real-time voice, video, and an avatar mode.
TypeSafe says jev-use generates candidate actions from browser state, has Jev select one, then validates and executes it through Cua Driver. Its CUA-S1-FORMS model scored forms locally in 7–9 ms, excluding execution.
CUA open-sourced CUA-S1-FORMS, a specialist model that selects bounded actions such as filling fields, checking boxes, clicking, or skipping. Cua Driver executes the ordered plan, and the MIT release includes synthetic-data generation, training, evaluation, and deployment tools.
Higgsfield showed GPT-6 Astra generating, editing, and publishing a video to TikTok. Its other demos have Astra operating Blender, After Effects, and Logic Pro for 3D scenes, animation, and audio work.
Higgsfield published demos of GPT-6 Astra controlling Figma, Photoshop, Procreate, After Effects, and Premiere Pro. The workflows show the model creating and editing visuals through those professional tools.
Higgsfield demos show GPT-6 Astra completing editable work in desktop creative software through its MCP integration. Examples include Blender scenes, After Effects motion graphics, Figma designs, DaVinci color grading, and Illustrator work.
Perplexity's Mac app can route sensitive agent steps to local models while using cloud models for other work. The company also open-sourced the PII classifier used to decide where work runs.
Anthropic added an isolated browser panel to Claude Cowork for navigating websites and completing forms. The rollout begins for paid desktop users, while Claude in Chrome is now generally available.
Perplexity’s Portable Computer runs its orchestrator, subagents, and harness locally on NVIDIA DGX Spark with a post-trained 27B model. Frontier-model escalation requires user approval and flags PII before text is sent externally.
OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.
Nous Research launched Hermes Agent with managed remote computers, provider and model choice, and local or hosted execution. Hermes Cloud idle instances start at 3 cents per day, according to the company.
Cua Driver's early-preview Computer History stores encrypted local metadata about agent actions so later sessions can recover successful routes. It excludes screenshots and typed text and is off by default.
Simon Willison quoted OpenClaw saying a gym API allowed cancelling other users’ reservations and moving a waitlisted user up one spot. Replies treated it as both an agent safety failure and a basic authorization bug.
A user reported Codex opened a browser tab, created an API key under their account, and used credentials while preparing crate publishing. The thread raised permission-boundary questions.
Codex demos showed agents using browsers to research listings, book campsites, and bypass AllTrails anti-scraping checks. Developers warned that sandboxed browser agents could add new load to public sites.
OpenAI said ChatGPT Work users can take over the cloud browser to sign in, then hand control back to the agent. The login persists across sessions, so repeat jobs can reuse the authenticated site without another takeover.
OpenAI staff said ChatGPT Work is available on web and mobile for Plus, Pro, Business, and Enterprise. Early users showed sites, email, documents, sheets, and browser tasks, while asking for memory and clearer docs.
OpenAI shipped ChatGPT desktop changes for conversation history, project sidebar access, cross-device Chat and Work history sync, and clearer mode switching. Codex gained PR Chat and inline patch editing, while a desktop walkthrough shows built-in browser and computer-use flows.
Perplexity released WANDR, its internal benchmark for deep and wide research in Computer. The dataset has 500 tasks, 170,495 source-backed records and production-derived use cases.
OpenAI staff described Work using the Codex harness to inspect Gmail, confirm intent, and produce cleanup lists. The rollout also added in-app code editing and a PR tab, while users questioned how Work, Chat, and Codex differ.
Practitioners ran GPT-5.6 Sol through Codex computer control on a five-hour Slay the Spire task and desktop fixes involving Chrome, 1Password, and a custom window utility. One report said Codex queued throwaway scripts for clicks and typing instead of driving every step from screenshots.
Perplexity added Grok 4.5 as an orchestrator model in Computer for Pro, Max, and Enterprise users. Perplexity reported a WANDR score of 0.328 at $4.76 per trial, while outside security-review and canvas-task tests put it close to GPT-5.6 Sol on cost or token use.
OpenAI launched ChatGPT Work, a Codex- and GPT-5.6-powered ChatGPT agent with a desktop app that can access local files and apps. OpenAI staff said Work and Codex use the same sandboxing with UI changes.
Meta launched Muse Spark 1.1 in Meta AI and the Meta Model API public preview for coding, tool use, computer use, and multimodal reasoning. Early eval posts ranked it highly while system-card threads flagged safety details.
OpenAI says GPT-5.6 Sol, Terra, and Luna will launch publicly Thursday as preview access expands. Testers describe Sol as strong for coding, agents, and computer use; Wafer reports Cerebras serving up to 750 tokens/sec.
CMU introduced Gym-Anything, which uses one agent to create software environments and another to audit screenshots, logs, files, and checklists. The project targets verified computer-use training tasks from ordinary desktop apps.
Anthropic launched Claude Sonnet 5 across Claude, the API, and Claude Code with 1M context, adaptive thinking, and $2/$10 intro pricing through Aug. 31. Independent evals place it near Opus 4.8 on coding and tool use, so teams should benchmark it against Opus before switching.
A day after Gemini 3.5 Flash Computer Use surfaced as a launch story, Google formally opened it through the Gemini API and Enterprise Agent Platform. Explicit user confirmation, automated task stopping, and an Android adb quickstart make the rollout concrete for agent builders.
Google released built-in Computer Use for Gemini 3.5 Flash across browser, mobile, and desktop. Try it for agent workflows, but watch for timeout issues on long design-from-scratch runs.
Hermes Agent added GUI computer-use support for Windows and Linux through TryCua drivers, extending beyond existing macOS support. Teams running desktop automation across mixed operating systems should test the new coverage.
OpenAI added Record & Replay to Codex so users can demonstrate a repetitive computer task once and save it as a reusable skill. The first rollout is Mac-only and unavailable in the EEA, UK, and Switzerland, so teams should check access before planning rollout.
TryCua brought Cua Driver to Linux, letting Claude Code, Codex, Hermes, and custom agents control real desktop apps via CLI or MCP without taking over the main terminal. The release also adds headless SSH execution and a preview of multi-window Wayland control across supported distros.
OpenAI expanded Codex in Europe with Computer Use, the Chrome extension, Memory, and Chronicle. The rollout broadens browser and desktop automation outside the U.S., though some memory features remain opt-in or preview-only.
ENPIRE launched a physical autoresearch setup that gives eight Codex agents robots, GPUs, and real-world APIs for tasks like zip ties and part sorting. It matters because it moves long-horizon agent evaluation from browser-only loops into embodied experimentation with explicit safety controls.
TryCua and Snorkel opened Cua-Bench, a computer-use benchmark with 25 expert-authored KiCad tasks graded by exact netlist matches. The early results show frontier models still struggle with GUI execution, wiring completion, and self-checking, so treat benchmark wins as incomplete for real computer-use work.
OpenAI shipped a docs agent that can hand off guides to Codex, and users published Appshots, browser-control, parallel PR, and multi-tree workflows. Watch the examples for ways to structure Codex around orchestrated tasks, while code-review and plugin gaps remain.
Perplexity made Deep Research a native skill inside Computer and tied it to the same harness, long-running sandboxes, tools, connectors, and licensed data. The update collapses multi-step research into one persistent agent interface instead of a separate mode.
Codex usage moved further into phone-first workflows, with iOS dictation loops, background voice capture, and app updates like searchable settings and restored state. The comparisons still flag rough spots in multi-thread UX, Windows support, and cases where CLI tabs or cloud agents are easier to manage.
Browser Use launched synced cloud profiles for logged-in sessions, added geo-targeted proxies, and showed a 484-browser startup demo that finished in under two seconds. The update matters because hosted browser agents can now keep authenticated state and regional routing without custom session-management work.
Perplexity opened Personal Computer for Windows to Max and Enterprise Max users on a waitlist. The rollout widens its local agent surface beyond earlier releases, and users should watch for the local-cloud task splitting preview for private or heavier workloads.
H Company released Holo 3.1, a local computer-use VLM family with function calling and AndroidWorld gains up to 79.3% on the 35B model. The update pushes computer-use agents toward local and mobile deployment instead of cloud-only runtimes.
OpenAI added computer use to Codex on Windows and lets ChatGPT mobile steer tasks running on Windows PCs. The update extends Codex to existing Windows dev machines and adds remote review and debugging from mobile.
Cua Driver said its Windows backend is now stable, letting Claude Code, Codex, Hermes, or custom agents drive real Windows apps through MCP or CLI. The release targets Windows-only line-of-business software while keeping the desktop usable with multi-pointer support.
Practitioners published reusable Codex workflows for project audits, memory-driven skill packaging, mobile delegation, and remote computer use. Try the prompt-and-steps patterns if you want to adapt Codex across repos and devices.
Two days after Codex added locked-Mac control and Appshots, users posted end-to-end iPhone simulator debugging, Safari form-filling, and remote-control workflows. That matters because the feature is moving from launch copy into concrete computer-use tasks that can replace manual QA and repetitive UI work.
OpenAI shipped a Codex update that lets the mobile app control a locked Mac, adds Appshots for screen context, and graduates /goal. It also adds browser annotation tools, team plugin sharing, and expanded analytics for business users.
Cognition added native Windows VMs to Devin so it can build, run, and test Windows applications with MSBuild, IIS, PowerShell, and SQL Server. The rollout lets Devin handle enterprise codebases where Linux sandboxes are not enough.
Leak videos and tester reports pointed to a larger Gemini desktop app with Stream to Cursor, Spark local-file access, Live, and Omni ahead of I/O. Independent testers also reported faster 3.2 and 3.5 Flash checkpoints, but Google had not announced the features publicly.
OpenAI documented Codex remote connections, letting the ChatGPT app point at a separate Codex host such as a Mac mini or rented VPS. Try it for long runs that need to stay alive off-device or for phone-first coding sessions.