Skip to content
AI Primer
TOPIC20 stories

KV Cache

Cache hit rate, offloading, routing, and cache-aware systems.

RELEASE25th September
OpenRouter launches Jev Router for cache-aware model routing

OpenRouter launched Jev Router, which selects a model and reasoning effort per turn while weighing the cost of losing cached context. OpenRouter reports 237 of 423 tasks solved; invalid Jev outputs or timeouts fail without fallback.

RELEASE24th September
Quail open-sources MIT-licensed AI-SQL engine for LLM queries

Quail open-sourced an MIT-licensed engine that plans AI queries, batches inference, and reuses KV cache across filters and joins. Its authors report 1.84× faster execution than hand-tuned vLLM on 29 queries.

WORKFLOW19th September
Teknium reports Jev compaction can increase token costs

Teknium’s public evaluation says a Jev compaction strategy removes tool calls and eventually stops yielding savings. Repeated compaction can invalidate caches and increase total token costs, according to the critique.

RELEASE2w ago
DeepSeek V4.1 Flash tops independent open-weight evaluations

DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.

WORKFLOW1mo ago
Speculative decoding tests report acceptance drop from 0.71 to 0.18 after ~32K context

A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.

NEWS1mo ago
Google DeepMind introduces SkillSmith for KV-cache skill composition in Gemma 3 4B

Google DeepMind introduced SkillSmith, a method that treats prefix weights or KV-cache states as an input modality so a frozen Gemma 3 4B model can synthesize new skill prefixes at inference time. Reported Composite-SNI Elo improved when cache composition was combined with text descriptions, making it a research artifact rather than a deployable runtime.

RELEASE1mo ago
Together releases ThunderAgent for KV-cache scheduling in agent workflows

Together released ThunderAgent to schedule whole agent workflows instead of isolated requests during tool calls. Together reports up to 2.5x throughput and about 10x lower P50 latency under high concurrency.

NEWS2mo ago
Baseten reports LLM fact writes can retain lift but vanish from answers

Baseten reports that injected facts can retain log-probability lift in weights while disappearing from answers, with prompts recovering 77–80% accuracy. OpenWiki adopted Open Knowledge Format for portable memory files.

NEWS2mo ago
Shepherd supports agent-run rollback beyond git rewind

A Stanford-linked Shepherd thread described rollback for long agent runs that restores files, packages, dev servers, and process state beyond git-style rewinds. Replies flagged KV-cache warmth and registered database inverse steps as cost and recovery boundaries.

RELEASE3mo ago
vLLM releases v0.22.0 with 28.9% FP8 latency cuts and KV offloading

vLLM 0.22.0 shipped DeepSeek V4 hardening, a Rust frontend, batch-invariant Cutlass FP8 paths, and multi-tier KV cache offloading. The release also removes deprecated APIs, so some serving stacks will need upgrade work.

RELEASE4mo ago
SGLang 0.5.12 adds DeepSeek V4 serving with ShadowRadix and HiSparse

SGLang v0.5.12 added native DeepSeek V4 support with ShadowRadix prefix caching, HiSparse CPU-extended KV, MegaMoE kernels, and Blackwell MLA work. The release broadens hardware targets and improves long-context serving efficiency for open runtimes.

RELEASE4mo ago
DeepSeek removes visual-primitives repo after 90-KV vision details

DeepSeek briefly published a paper and threads on point-and-bbox reasoning, about 90 KV entries per 800² image, and RL-trained vision experts, then removed the repo and related mentions. The technique looked like a low-token path to computer use and multimodal reasoning in V4-Flash, but availability and reproducibility are now unclear.

RELEASE5mo ago
vLLM 0.20.0 releases TurboQuant 2-bit KV cache, CUDA 13 baseline, and DeepSeek V4 upgrades

vLLM 0.20.0 shipped a new CUDA 13 / PyTorch 2.11 / Transformers v5 baseline, TurboQuant 2-bit KV cache, FA4 MLA defaults, and deeper DeepSeek V4 support. The release changes serving baselines across NVIDIA, AMD, Intel, and ARM-CUDA setups, including 4x KV capacity and a clearer upgrade path for teams already running V4.

NEWS5mo ago
Moonshot claims 1.54x throughput and 64% lower P90 TTFT with cross-datacenter prefill

Moonshot says its Prefill-as-a-Service setup makes prefill/decode disaggregation practical across datacenters and mixed hardware by shrinking KV cache with Kimi Linear. The paper reports 1.54x throughput and a 64% drop in P90 time-to-first-token, so benchmark the approach before planning production adoption.

NEWS6mo ago
TurboQuant updates 2.5-bit mixed precision with PyTorch and llama.cpp ports

New discussion around TurboQuant focuses on its 2.5-bit mixed-precision setup and working PyTorch and llama.cpp implementations. The technique is moving from a research claim into deployable KV-cache compression with concrete porting details.

NEWS6mo ago
TurboQuant cuts KV cache memory 6x with 3-bit storage

Google Research said TurboQuant can shrink KV cache storage to 3 bits with roughly 6x less memory, and early implementations already surfaced in PyTorch, llama.cpp, and Atomic Chat. The work targets a core inference bottleneck for long-context serving on local and server hardware.

NEWS6mo ago
Google Research launches TurboQuant: 6x KV-cache compression, 8x faster H100 attention

TurboQuant claims 6x KV-cache memory reduction and up to 8x faster attention on H100s without retraining or quality loss on long-context tasks. If those results hold in serving stacks, teams should revisit long-context cost, capacity, and vector-search design.

NEWS6mo ago
Flash-MoE claims Qwen3.5-397B runs on iPhone at 0.6 tokens/sec via SSD streaming

Flash-MoE now shows SSD-streamed expert weights pushing a 397B Qwen3.5 variant onto an iPhone at 0.6 tokens per second, extending its earlier laptop demos. Treat it as a memory-tiering prototype rather than a deployable mobile serving target, because speed, heat, and context headroom remain tight.

RELEASE6mo ago
H Company releases Holotron-12B: 8.9k tok/s on H100 and 80.5% WebVoyager

H Company launched Holotron-12B, an open multimodal model for computer-use agents built on a hybrid SSM-attention stack that targets KV-cache bottlenecks. Benchmark it if you need high-concurrency browser agents and want better throughput without giving up web-task accuracy.

WORKFLOW6mo ago
oMLX supports Claude Code locally with tiered KV cache and Anthropic Messages API

oMLX now supports local Claude Code setups on Apple Silicon with tiered KV cache and an Anthropic Messages API-compatible endpoint, with one setup reporting roughly 10x faster performance than mlx_lm-style serving. If you want private on-device coding agents, point Claude Code at a local compatible endpoint and disable the attribution header to preserve cache reuse.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.