CloudMatrix-Infer
Comprehensive LLM serving solution for Huawei CloudMatrix384
CloudMatrix-Infer is a comprehensive LLM serving solution proposed for Huawei CloudMatrix384. It is designed for large-scale MoE model inference such as DeepSeek-R1, using peer-to-peer serving, disaggregated prefill/decode/caching resource pools, large-scale expert parallelism, and hardware-aware optimizations on Huawei Ascend/CloudMatrix infrastructure.

Recent stories
0 linked stories
No linked stories yet.