Skip to content
AI Primer
workflow

Firstmate cuts AGENTS.md instruction file by nearly 50%

A case study reports reducing a roughly 30,000-token AGENTS.md file by almost 50% while improving instruction quality. Session transcripts informed which rules remained and which moved into a skill.

4 min read
Firstmate cuts AGENTS.md instruction file by nearly 50%
Firstmate cuts AGENTS.md instruction file by nearly 50%

TL;DR

  • Firstmate nearly halved its startup instructions: kunchenguid's case study reports a 48% reduction after analyzing 100 session transcripts.
  • Most savings came from moving conditional instructions into on-demand skills, using the distinction between constantly beneficial and situational rules described in a reply.
  • The edits added missing guidance and tightened existing rules, while an eval showed no regression, according to a follow-up.

Firstmate's token counts use a bytes/3 estimate, and the PR records follow-up fixes for gaps in when skills load. Backpass's best detail is its rule for missed skills: it proposes editing the description when an agent overlooks knowledge already inside the skill.

30k to 16k estimated tokens

Firstmate is an agent distro whose supervisor dispatches coding agents into separate git worktrees and collects finished PRs. Its AGENTS.md defines core supervisor behaviors and loads at the start of every main Firstmate session.

Kun Chen, backpass's creator, had already manually pruned that file to roughly 30,000 tokens. He stopped because further cuts caused regressions.

With backpass, Chen analyzed 100 recent session transcripts, reviewed the proposed changes, and rejected two edits over their rationale and return on investment. The file ended up at about 16,000 estimated tokens in the merged PR, which reports a 47% reduction; his post reports 48%.

Seven on-demand skills

Firstmate extracted seven situational sections into agent-only skills that load when their triggers fire:

  • operational-home-layout: home, configuration, data, state, and project paths.
  • session-start-recovery: unfinished checks, diagnostics, and recovery inputs from the session digest.
  • validation-supervision: active validation runs, mid-run requirement changes, and findings.
  • ship-landing: PR readiness signals, landing, and cleanup.
  • scout-completion: scout completion, visual iteration, and promotion.
  • away-quiet-supervision: /afk and /quiet behavior.
  • agent-skill-trigger-index: the complete agent-only trigger index for audits.

Relay activation and ownership rules also moved into the existing fmx-respond skill.

Backpass evidence gates

Backpass processes transcripts through five stages:

  1. Distill: preserve user and assistant turns, collapse tool calls, truncate tool output, and redact obvious secrets.
  2. Analyze: make a cheap model call per transcript to identify helpful instructions, violations, and uncovered mistakes. Claims without verbatim quotes are discarded.
  3. Aggregate: calculate per-instruction counts and relevance, measured as the share of sessions in which an instruction mattered. New instructions require evidence from at least two independent sessions.
  4. Propose: use a high-reasoning session to produce ADD, REMOVE, REWRITE, or EXTRACT→SKILL edits in a staging copy.
  5. Review: show each edit and its evidence through backpass apply; analysis does not modify the target memory file.

Its routing policy keeps safety-critical instructions in always-loaded memory regardless of frequency. Conditional instructions need a detectable trigger to become skills.

The 100-session sample can be expanded at additional analysis-token cost. Chen also said human reviewers judge which omissions would be “costly when missed.”

Explicit Bash for production checks

Firstmate's coding-guideline update added explicit Bash execution for two cases:

  • Tests of the production library.
  • Commands that source scripts under bin/.

Regression checks

Chen said his private eval set stayed isolated from the optimization data to avoid overfitting.

A separate live comparison of the branch and main used fresh Claude-backed Firstmate sessions. The PR reports no core-behavior regression, but a follow-up commit repaired load-timing gaps and references:

  • Load validation-supervision on every ask-user answer.
  • Keep guidance for mid-task captain asks inline.
  • Keep unconfirmed network-check guidance inline.
  • Keep worker account-pin rules inline.
  • Correct stale cross-references.

Handwritten user preferences

Chen pointed a reader back to his August 23 blog post, where he assigns the two memory levels different treatment:

  • User-level memory: small, handwritten preference files such as ~/.claude/CLAUDE.md or ~/AGENTS.md. In his setup, tools and agents do not optimize these files.
  • Project-level memory: repository instructions refined from session evidence within a token budget.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 1 thread
Handwritten user preferences1 post
Share on X