LMCache
Supercharge Your LLM with the Fastest KV Cache Layer
LMCache is an open-source KV cache management layer for scalable LLM inference. It stores, reuses, searches, compresses, moves, and observes LLM KV caches across serving engines and storage backends to reduce time-to-first-token and improve throughput for long-context, multi-turn, agentic, and RAG workloads.

Recent stories
0 linked stories
No linked stories yet.