Skip to content
AI Primer

Mooncake

A KVCache-centric Disaggregated Architecture for LLM Serving.

Mooncake is an open-source KVCache-centric disaggregated serving platform and infrastructure for LLM inference. It separates prefill and decode clusters, uses underutilized CPU/DRAM/SSD resources to build a disaggregated KV cache pool, and provides components such as Mooncake Transfer Engine and Mooncake Store for high-performance KV cache transfer, storage, and reuse across serving instances.

Screenshot of Mooncake website

Recent stories

0 linked stories
No linked stories yet.
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.