Sherpa scores 89.8% on 176 cross-chapter memory questions, Pocket FM says
Pocket FM launched Sherpa, a beta system that uses planner, feedback, and storyboarding agents to produce serialized fiction. Pocket FM reports 89.8% accuracy across 176 cross-chapter story-memory questions versus 57.4% for a graph-memory baseline.

TL;DR
- Sherpa is Pocket FM’s free beta writing partner for serialized fiction, dividing work among planning, editing feedback, and a persistent story record, as testingcatalog’s launch post reports.
- Pocket FM reported 89.8% accuracy on 176 cross-chapter questions, versus 57.4% for Graphiti, according to testingcatalog’s metric summary.
- The claimed memory lift comes from typed narrative state, including who knows what, event versus reveal order, and setup-to-payoff status, as kimmonismus’ explainer describes.
The Narrative World Model paper fixes the answering model at Opus 4.8, leaving each memory system to supply its own chapter-safe evidence. Pocket FM’s product page also claims a separate 1,204-continuation evaluation across 38 series, with tests reaching 200 episodes.
Narrative World Model
Pocket FM publishes finalized chapter prose into a state system, then retrieves only material available at the selected chapter. The paper says a later rewrite invalidates downstream graph edges from the revised chapter.
Its writer-facing records include:
- character knowledge and unknowns
- event order distinct from reveal order
- relationship changes and object state
- open or closed promises and payoffs
- focalized observer and dramatic function
The retrieval path combines BM25, vector search, and one-hop expansion over a temporal knowledge graph. Every answer is constrained to source chapters at or before the requested checkpoint.
Planner, Feedback, Storyboard
Sherpa exposes three separate agents, with the product page filling in their operating roles:
- Planner maps arcs, character journeys, conflicts, and episode beats against the full story context.
- Feedback evaluates engagement, readability, prose, coherence, pacing, and enjoyment, then proposes revisions.
- Storyboard tracks characters, settings, arcs, scenes, and decisions across drafts and episodes.
The Narrative World Model sits beneath that split as the shared source of story state, rather than treating each generation as an isolated prompt.
The 176-question test
The 89.8% result is from a private, curated set of 176 questions that each require evidence from at least two chapters. Pocket FM’s paper says the set spans four question families: character knowledge, event-versus-reveal order, setup/payoff structure, and combinations of those forms.
The comparison held the reader constant at Opus 4.8, while NWM Graph Retrieval and Graphiti supplied different chapter-filtered evidence. The same paper reports a public 576-question result of 62.5% for NWM and 51.6% for Graphiti, figures also cited by kimmonismus.
The protocol isolates memory and retrieval behavior from text generation. Its public corpus contains 12 public-domain books across six genres, while the five 50-chapter production-style serials behind the private test are not redistributable; the document is under review as a COLM 2026 paper.
Localization and publishing
Pocket FM says Sherpa can localize a draft into seven languages, narrate it with a consistent voice for a season, and publish it as a Pocket FM episode. The company’s audio-drama generator page says its localization layer adapts names, places, idioms, and cultural references, then sends the narrated chapter directly into its publishing flow.