Skip to content
AI Primer
workflow

Report: Asana cuts browser-agent cost from $1.97 to $0.47 per run with prompt caching

Asana reportedly reduced its GPT-6.1 Sol browser-agent cost from $1.97 to $0.47 per run with an 89% cache hit rate. The report describes append-only prompt history and a Codex cache-measurement command.

6 min read
Report: Asana cuts browser-agent cost from $1.97 to $0.47 per run with prompt caching
Report: Asana cuts browser-agent cost from $1.97 to $0.47 per run with prompt caching

TL;DR

  • Asana’s GPT-6.1 Sol runs fell from $1.97 to $0.47 in estimated model cost, with 89% of input served from cache, as daniel_mac8 reported.
  • The winning configuration cached browsing history, retained more text, and removed screenshots in batches, according to the linked case study.
  • Long sessions encounter a pricing boundary above 272K input tokens; daniel_mac8’s follow-up raised compaction as the tradeoff.
  • Codex’s Steer setting accepts follow-ups during a running task, as demonstrated in the desktop demo.

The winning screenshot policy accumulated 20 screenshots, then kept only the newest one, preserving history across roughly 19 consecutive calls. Sol’s cache writes cost 1.25 times ordinary input, while cache reads cost just 5% of that ordinary rate.

Cost per run

Asana’s most useful result held the model constant: with the larger history budget, changing GPT-6.1 Sol’s caching and screenshot policy cut estimated model cost from $1.97 to $0.47 per run. In OpenAI’s October 9 case study, 89% of input tokens came from cache, and each call was about three times cheaper.

The separate 76x comparison starts with the original production setup on Model B. Its $36.21 baseline is a lower bound because some runs hit the step limit before finishing; runtime fell from at least 22.5 minutes on that setup to roughly four minutes with optimized Sol.

20-to-1 screenshot pruning

The original agent cached its instructions and tool definitions but repeatedly paid for its growing browsing history. Removing the previous screenshot at every step and trimming older text changed the prefix that later calls needed to reuse.

Asana describes three fixes in its engineering write-up:

  1. Cache the history: place a cache marker on the latest tool result.
  2. Retain more text: increase the history budget from 120,000 to 480,000 characters.
  3. Prune screenshots in batches: accumulate up to 20, then cut back to the latest one.

Caching history without changing pruning actually increased cost on three of four models at the larger budget. The agent kept rewriting the cache and rarely reading it.

144 runs on a book catalog

Asana used GPT-6 Astra in Codex to audit request construction, instrument calls, refactor the code for parallel workflows, and analyze traces. Frank Hidalgo, StackAI CTO at Asana, estimated that the work took about a week rather than one to two months by hand.

The experiment had a concrete, narrow workload:

  • Four models: GPT-6.1 Sol and anonymized Models A, B, and C.
  • Twelve conditions per model: six caching and screenshot policies at two history budgets.
  • Three repetitions per condition: 144 main-study runs.
  • One task: collect six fields for each of 32 books from a public demo catalog, totaling 192 facts.
  • Scoring: costs calculated from provider token counters; answers checked against an independently prepared reference.

Requests, traces, and results were recorded in Command, Asana’s software delivery platform. Findings became tickets and reviewed pull requests, and the browser-navigation changes shipped to StackAI.

History budgets and answer completion

More history also changed whether the agent returned an answer. Across the six policies, Asana reported these results in its study:

| Model | 120,000-character budget | 480,000-character budget |
| --- | ---: | ---: |
| GPT-6.1 Sol | 3 of 18 runs answered | 18 of 18 answered |
| Model C | 0 of 18 runs answered | 18 of 18 answered |

All 18 larger-budget Sol answers were correct, according to OpenAI’s account. No run reached the 480,000-character cap, so that limit never constrained these experiments.

Four Codex cache habits

The four habits in daniel_mac8’s post translate into four request-level mechanics:

  1. Append-only sessions: preserve earlier messages and add new context at the end, rather than rewriting history or restarting the session.
  2. Stable configuration: model, tool, and reasoning-effort changes can alter the beginning of a request. The post identifies GPT-6 Astra as an effort-change exception.
  3. Compaction boundaries: /compact rewrites earlier context, changing the reusable prefix from the first altered token onward.
  4. Token accounting: cached-input fraction is cached_input_tokens / input_tokens.

The command in the post is:

OpenAI’s non-interactive CLI documentation specifies a JSON Lines stream and places those usage fields on the turn.completed event.

Append-only configuration updates

OpenAI documents two ways to change behavior while preserving earlier context in its caching guide:

  • Reasoning effort: supported GPT-6 models accept an appended configuration_update item while the top-level reasoning.effort stays at its original value.
  • Tool availability: allowed_tools restricts callable tools without replacing their definitions; tool_choice: "none" disables calls while leaving those definitions intact.

The documented effort update is:

In another reply, daniel_mac8 also claimed that Claude Opus 5.5 preserves its cache across reasoning-effort changes.

Compaction at 272K tokens

Sol’s long-context pricing starts above 272,000 input tokens. The published API prices, per million tokens, are:

| Token category | Short context | Long context |
| --- | ---: | ---: |
| Input | $2.00 | $4.00 |
| Cached input | $0.10 | $0.20 |
| Cache writes | $2.50 | $5.00 |
| Output | $10.00 | $15.00 |

Daniel_mac8 argued for autocompaction before that boundary, adding that he believed Codex already used it as the default. Compaction itself was outside Asana’s study.

Real-time steering

Codex follow-ups have two behaviors in OpenAI’s desktop settings documentation:

  • Queue: wait for the current run to finish before sending the follow-up.
  • Steer: apply guidance to the work already in progress.

Daniel_mac8 demonstrated Settings → General → Follow-up behavior → Steer and described steering as real-time.

Keeping every screenshot

Asana then ran a 12-run follow-up that retained every screenshot. Cost per call was 1.2x lower than the best batch-pruning condition on Model B and GPT-6.1 Sol, and about 5% lower on Model C.

The authors describe the results as broad patterns: with three or four runs per condition and varying call counts, the study cannot reliably distinguish configurations separated by only a few percent.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR1 post
Cost per run1 post
Four Codex cache habits1 post
Append-only configuration updates2 posts
Compaction at 272K tokens1 post
Share on X