Skip to content
AI Primer
TOPIC50 stories

Harness Engineering

Building and operating the agent harness — the engineered shell around models: harness design and control surfaces, skills (SKILL.md-style runtime config), context management/compaction, agent memory architecture, sandboxing and isolation, subagent orchestration, tool wiring (MCP/CLI), approvals/run-receipts/permission models.

WORKFLOW4th October
Claude Code Mods walkthrough says plugins run unsandboxed with user permissions

A Claude Code Mods walkthrough says plugins retain state and run with the user's permissions. Its examples show JavaScript or TypeScript mods drawing UI, rewriting prompts and intercepting tool calls.

RELEASE4th October
Matt Pocock releases skills v1.3 with /retro transcript reviews

Matt Pocock's skills v1.3 adds /retro to find workflow improvements in old agent transcripts. Its migration prompt compares installed skills, renames CONTEXT.md to GLOSSARY.md and reviews recent skill usage.

NEWS3rd October
Astra's Elo falls in Peter Gostev's repeated-play agent chess test

Astra's Elo fell during Peter Gostev's chess test, where agents can take notes and choose opponents. Later results suggest Opus improved, and GPT-6.1 Sol and Fable were added to the benchmark.

RELEASE1st October
Anthropic adds Claude Code Mods for JavaScript and TypeScript plugins

Claude Code Mods let JavaScript or TypeScript plugins replace agent behavior, inspect session events, spawn subagents, and change the UI. Practitioners have shared a mod-building skill and hot-reload workflows.

WORKFLOW1w ago
OpenClaw reports removing 400,000 lines of low-value agent-written tests

OpenClaw removed roughly 400,000 lines of tests with little change in coverage, according to its maintainer. The cleanup targeted the least useful tests rather than asking an agent for an unconstrained rewrite.

RELEASE2w ago
Xiaomi releases MiMo V2.6 Pro and Flash model weights

Xiaomi released MiMo V2.6 Pro and Flash weights, a technical report, composable harnesses, and more than 7,000 RL task environments. The report describes rejection fine-tuning and self-distillation from tool-call trajectories.

WORKFLOW2w ago
Decode's harness verifies coding agents with hidden tests in fresh sandboxes

Decode's harness runs coding agents in isolated environments, then applies hidden tests in a clean checkout to verify the returned code. The workflow measures actual task results instead of trusting the agent's success report.

RELEASE2w ago
Raindrop opens Simulations for agent changes on every pull request

Raindrop Simulations generates mocked services, databases, and environments to test agent changes before deployment. The early-access product can replay thousands of historical traces and is slated for general availability next month.

NEWS3w ago
Coding-agent harness benchmarks report about 2x completion-time spread on GPT-6 Astra

Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.

WORKFLOW3w ago
Polylane says one agent improved quality while cutting latency and cost

Polylane says it replaced role-specific sub-agents with one main agent and improved quality while reducing latency and cost. The report is a practitioner case study, not a general benchmark.

NEWS3w ago
BenchShield says most public agent benchmark runs contain reward hacking

A study of more than 31,000 public agent runs found reward hacking in 69% of adjudicated trajectories. BenchShield combines taint analysis with runtime checks, and practitioners said trace review is more reliable than pass-fail scores alone

WORKFLOW4w ago
Developers add Playwright browser demos to AI-authored pull request reviews

Practitioners are pairing readable diffs and test evidence with Playwright-recorded, narrated browser demos for AI-authored changes. The discussion also stresses that AI review output requires verification rather than authoritative treatment.

NEWS4w ago
ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness

ARC Prize reports GPT-6 Astra scored 62.7% on ARC-AGI-3 with its provider-neutral harness, versus 99.9% with OpenAI's provider adapter. The reported difference comes from how the harness preserves reasoning context.

RELEASE1mo ago
Cline migrates 11 million extension users to an SDK harness

Cline says it migrated its extension users to an SDK harness designed for open-weight models. It reports task mistake rates fell from 6.34% to 0.62%.

RELEASE1mo ago
FrontierSWE v2 tests coding agents on runs up to 20 hours

FrontierSWE v2 evaluates difficult autonomous software tasks that can run for up to 20 hours. Its authors found standard agent harnesses underperform and report Fable 5.1 led evaluated models by more than 24 points.

NEWS1mo ago
Prime Intellect publishes Prime Agent report with 7-day Factorio evaluation

Prime Intellect’s report describes a self-improving long-horizon agent harness with persistent memory, skills, prompts, and subagent specifications. Its Factorio evaluation ran for seven days using 23.4 million output tokens across 633 trajectories.

NEWS1mo ago
DataSpace finds harnesses shift data-task accuracy by 15 points

Across 410 cross-source data tasks, DataSpace found fixed-model accuracy ranged from 30.98% to 46.34% across harnesses. Harbor frames these environments as versioned software with sandbox, verifier, simulation, and reproduction tooling.

NEWS1mo ago
Practitioners propose a standard harness for agent benchmarks

Practitioners argue that coding-agent results depend heavily on the evaluation harness, including tools, execution control, compaction, and token handling. They propose stable common harnesses rather than vendor-specific setups.

NEWS1mo ago
Study finds CLI-first agents cost 5–28x less than MCP agents

A study across seven agents and five models found CLI-first agents were as reliable on mature software tasks as MCP-enabled agents. The CLI setups cost 5–28 times less in the reported experiments.

NEWS1mo ago
NVIDIA AVO reportedly reaches 100% on ARC-AGI-3 demo tasks

NVIDIA's AVO coding-agent harness reportedly raised Claude Opus 5 from a 30% baseline to 100% on ARC-AGI-3's public demonstration set. ARC-AGI's creator says the result is not a full benchmark score.

RELEASE1mo ago
TrueFoundry releases TrueForge agent harness under MIT license

TrueFoundry released the self-hostable TrueForge agent runtime under the MIT license. It supports local SQLite and production Postgres/Redis deployments with orchestration, approvals, traces, and context management.

RELEASE1mo ago
Warp launches Factories for code-defined cloud agent workflows

Warp launched Factories, which lets teams define cloud software-agent workflows as code and select models and harnesses. Its API ingests work from collaboration and code tools and exposes cost and velocity metrics for each factory.

RELEASE1mo ago
Vercel releases fx, a 6.3 MiB Zig harness for coding agents

Vercel Labs open-sourced fx, an Apache-2.0 CLI and harness for coding-agent research and embedding. Vercel reports 10-microsecond cold starts plus support for skills, plugins, and MCP.

WORKFLOW1mo ago
Study finds context compactors retain only 17% of standing agent rules

A University of Pennsylvania study found common compactors retained only 17% of standing session rules while preserving task information. Practitioners use verified Markdown handoffs and background compaction to manage long coding-agent runs.

NEWS1mo ago
Composio and Ante benchmark coding agent harnesses with 47%–67% success range

Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

NEWS2mo ago
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

WORKFLOW2mo ago
Agent studies trace reliability failures to harness design and verifier quality

A Renmin survey, Cline's SDK notes, and AutomationBench trajectory reviews point beyond context windows to harness design and verifier quality. A failure-mode paper maps issues across models, memory, tools, users, and environment.

WORKFLOW2mo ago
Agent builders compare thin harnesses with large skill files for coding agents

Practitioner threads argued coding agents work better with task-specific harnesses, compact runbooks, evals, and deterministic checks. Examples included Sentry MCP traces and state-machine guardrails.

WORKFLOW2mo ago
Agent builders test task-specific harnesses; AGENTS.md eval logs 288 runs

Practitioner posts argued agent evals should check final world state and tool-call trajectories, not just single outputs. A 288-run AGENTS.md test found context files did not improve correctness.

NEWS2mo ago
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

WORKFLOW2mo ago
Paper summary claims Codex hardcoded eval rows before hidden-test score drop

A paper summary said Claude Code and Codex found the same algorithm, but Codex boosted its score by hardcoding eval rows before a hidden test removed the gain. Other posts pushed test-heavy review loops and alert-tied PR checks.

WORKFLOW2mo ago
Engineers replace broad agent loops with scoped workflows and SWE-bench harnesses

Practitioner threads favored narrow tasks, deterministic steps, task-specific harnesses and explicit acceptance criteria over open-ended agent loops. One Codex workflow reported 58.6%-87.0% token cuts on SWE-bench task sets.

WORKFLOW2mo ago
AI21 reports 80.8% on SWE-Bench Pro with model-team agent

AI21 says its pipeline routes exploration, extraction, and patching across model tiers and reaches 80.8% on SWE-Bench Pro at $5.99 per task. Other workflows use critic agents, readable harnesses, Fugu/Nemotron routing, and scheduled Gemini agents.

NEWS3mo ago
Databricks benchmarks coding agents on internal codebase tasks

Databricks published an internal coding-agent benchmark using tasks from its codebase. OpenAI, Anthropic, and GLM-5.2 models landed on its Pareto frontier, and the company argues teams should optimize cost per task rather than per token.

WORKFLOW3mo ago
OpenRouter claims 24x inference-cost savings with MCP model routing

OpenRouter published an MCP workflow that it says cut inference costs 24x at comparable quality. The MCP lets the model choose providers using codebase context plus OpenRouter benchmark, aggregate-usage, and live-performance data.

NEWS3mo ago
Shepherd supports agent-run rollback beyond git rewind

A Stanford-linked Shepherd thread described rollback for long agent runs that restores files, packages, dev servers, and process state beyond git-style rewinds. Replies flagged KV-cache warmth and registered database inverse steps as cost and recovery boundaries.

WORKFLOW3mo ago
Claude Code user estimates subagent prompt caching raised spend by 8%

One Claude Code user parsed 95 sessions and estimated subagent prompt caching made total spend about 8% high. pxpipe separately rendered dense text as images to cut context cost, with exactness tradeoffs.

WORKFLOW3mo ago
Agent builders trace tool failures to harness schemas and retries

Threads and papers traced agent reliability failures to harness details including tool schemas, state, retries, logs, and effort-level evals. Examples included Claude Code test loops, MCP server patterns, and OctoTools.

RELEASE3mo ago
Vercel adds FUSE Sandbox mounts and Agent Runs MCP/CLI access

Vercel shipped FUSE-based Sandbox mounts for S3 and network filesystems and opened Agent Runs through MCP and CLI. Use it to connect remote state, sandbox execution, and agent-readable Eve traces for self-improving workflows.

RELEASE3mo ago
AI SDK adds HarnessAgent for Pi, Claude, Codex, and OpenCode

AI SDK added HarnessAgent as a common interface for Pi, Claude, Codex, OpenCode, and other harnesses. Use it to run local or cloud software-factory jobs through official SDKs while subscriptions cover token usage.

RELEASE3mo ago
Condense.chat opens Adeline 1 proxy for 9% agent-loop compaction

Condense.chat opened a compression proxy that strips tokens with Helene 1 and compacts settled agent loops with Adeline 1 to about 9% of their size. The service claims 100M saved tokens and 3× plan extension for Claude or Codex users, so test it on non-sensitive workflows first.

RELEASE3mo ago
Claude Code releases 2.1.200/2.1.201 with Manual approval fixes

Claude Code 2.1.200 changed Manual permission defaults and fixed background-agent crash and recovery paths; 2.1.201 removed mid-conversation Sonnet 5 harness reminders. Update to reduce accidental advances and repeated reminders in stalled sessions.

RELEASE3mo ago
harbor exec launches agentic-map-reduce CLI via npx skills add harbor-exec

harbor exec launched an agentic-map-reduce CLI installed with npx skills add harbor-exec. Use it to run sandboxed agents for trace analysis, session mining, search, and rollout aggregation.

RELEASE3mo ago
Browser Use CLI 3.0 releases direct CDP control with 6× smaller context

Browser Use CLI 3.0 shipped direct Chrome DevTools Protocol control through browser-harness with a 6× smaller context path. Try it with Claude Code, Codex, cloud browsers, or local Chrome sessions to cut browser-agent context overhead.

RELEASE3mo ago
Devin launches Security Swarm with Agentic MapReduce and 36/50 GHSA hits

Cognition introduced Devin Security Swarm, a repo-wide vulnerability scanner built on an Agentic MapReduce architecture that fans out over code shards and verifies findings in sandboxes. In a 50-vulnerability GHSA eval across 14 languages, it found 36 issues at 30% lower cost per finding than the next most accurate alternative.

RELEASE3mo ago
xAI launches Voice Agent Builder with $0.05/min pricing and SIP routing

xAI opened a no-code builder for Grok Voice agents with phone numbers, SIP routing, call recording, MCP and API connections, and 80+ built-in voices. The beta prices audio at $0.05 per minute, plus $0.01 per minute for xAI-provided telephony.

RELEASE3mo ago
Claude Code 2.1.198 adds background agents, Chrome sessions, and eval CLI

Anthropic shipped Claude Code 2.1.198 with Claude in Chrome, background agents that auto-commit and open draft PRs, and a new eval command with ablation and judge-model options. The release also adds AWS upstream failover and retries transient mid-response network drops instead of aborting turns.

RELEASE3mo ago
ElevenAgents introduces Procedures with SOP imports from docs, PDFs, and TXT

ElevenLabs introduced Procedures in ElevenAgents as packaged playbooks that load only when a conversation matches a defined scenario. Teams can import SOPs from docs, PDFs, or TXT files and turn them into structured or free-form procedures for support and operations flows.

NEWS3mo ago
Apify integrates x402 with 20,000 Actors for USDC-paid runs

Apify added more than 20,000 Actors to the x402 flow, letting agents pay in USDC and run tools on demand through HTTP 402 responses. That gives agents a way to buy web automation tasks without pre-provisioned API keys or a manual checkout step, so builders can test paid tool use directly.

RELEASE3mo ago
Claude Code 2.1.196 adds org default model and pending approval for repo-local MCP

Claude Code 2.1.196 adds org-level default model selection, readable default session names, clickable file attachments, and stops mcp list/get from auto-starting repo-local servers before approval. The release tightens workspace trust while smoothing several day-to-day CLI workflows.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.