Mooncake
A KVCache-centric Disaggregated Architecture for LLM Serving.
Mooncake is an open-source KVCache-centric disaggregated serving platform and infrastructure for LLM inference. It separates prefill and decode clusters, uses underutilized CPU/DRAM/SSD resources to build a disaggregated KV cache pool, and provides components such as Mooncake Transfer Engine and Mooncake Store for high-performance KV cache transfer, storage, and reuse across serving instances.

Recent stories
0 linked stories
No linked stories yet.