vix.ing · top · new · best · stats · spec

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

2024/06/24 by Ruoyu Qin, Qin, Ruoyu, Zheming Li +11 · 3 voices · 33 citations
Engineering · #Plasma Diagnostics and Applications #Radiation Effects in Electronics #Semiconductor materials and devices #cs.AI #cs.AR #cs.DC

paper · pdf · doi:10.48550/arxiv.2407.00079

openalex publication_date 2024/06/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.

Cited by

Discussions

Related