Claude Code user finds longer subagent caching saved no money
Daniel Mac shared a script to compare Claude Code subagent cache lifetimes, then reverted to five minutes after finding no savings for his workload. He also corrected the hour-long setting from 60 to 1h.

TL;DR
- Longer caching saved no money for Daniel Mac's workload; he reported switching back to five minutes after estimating the costs.
- The hour-long setting is
"subagentPromptCacheTtl": "1h", as his correction clarified after the original tip incorrectly used60. - Avoided cache rewrites have to offset more expensive writes; his script estimates that tradeoff from local Claude Code transcripts.
A cache hit resets the expiry timer without another cache-write charge. Mac's Python script counts the gaps where an hour-long cache could prevent a rewrite. The Claude Code docs give main conversations and subagents different default lifetimes.
Five-minute reversal
Daniel Mac initially warned that orchestrated subagents could lose cache pricing after sitting idle for more than five minutes. His refreshingly blunt follow-up reported the opposite result for his own workload: “In my case, the verdict is no. I switched it back to 5m.”
He also relayed feedback that the defaults worked for most people. His reported result came from a cost estimate, rather than a published comparison of bills before and after the change.
Corrected setting
Mac corrected the original screenshot's numeric 60 to the hour-long value:
Subagent idle gaps
The five-minute subagent default behind his exchange with doodlestein differs from the main conversation's default on a subscription. According to Claude Code's documentation:
- Main conversation within included subscription usage: one hour.
- Subagents: five minutes, including on a subscription.
- Main conversation using usage credits, an API key, or a cloud provider: five minutes.
Mac described orchestration patterns that can leave agents waiting:
- Passing work between parallel agents.
- Escalating work to a coordinator.
- Waiting for another agent to finish.
He distinguished those “digital worker” workflows from interactive coding in another reply.
Cache write pricing
Anthropic's prompt-caching pricing table expresses the charges as multiples of the model's base input-token price:
- Five-minute write: 1.25×.
- One-hour write: 2×.
- Cache read: normally 0.1×, with model-specific exceptions.
The one-hour write costs 60% more than the five-minute write. At the script's default read rate, each token newly written with the longer lifetime adds 0.75 base-input units, while replacing an expired five-minute rewrite with a cache read saves 1.15 units.
A longer lifetime generates savings only where it preserves a matching prefix that would otherwise need rewriting. Requests close enough together already keep the five-minute cache warm.
cache_gap_check.py
The published script uses only Python's standard library. It makes no network calls and writes no files.
Its mechanics are explicit:
- Scan
~/.claude/projects/**/*.jsonl, looking back 14 days by default. - Identify subagent records through
isSidechain,agentId, or asubagentsdirectory, excluding ordinary main-thread records. - Deduplicate streamed records by request or message ID, retaining the earliest timestamp and later usage data.
- Count consecutive-request gaps below five minutes, from five to under 60 minutes, and at least 60 minutes.
- Compare recorded cache costs with hypothetical one-hour costs, expressed in base-input-token units.
The verdict labels are worth-it for estimated savings above 5%, not-worth-it for an increase above 5%, and marginal within that band. The baseline preserves any already-recorded one-hour writes at their higher rate.
Estimator assumptions
The estimator projects extra cache hits using timestamps and token counts, without comparing prompt contents. It assumes the preceding cached prefix remains reusable across a five-to-60-minute gap; Anthropic's exact-prefix matching requirement can make that assumption optimistic when the prompt changes.
The modeled totals include cache reads and writes, excluding uncached input and output charges. Its default read multiplier is 0.1, configurable through --read-mult; Anthropic's pricing documentation also lists 0.05× and 0.025× rates for some models.