Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
C++kvcache-ai/Mooncake

Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

87.5/100
6.5KForks: 1.2K
View on GitHubHomepage →
Loading report...

Similar Projects

vllm-ascend

79

Community maintained hardware plugin for vLLM on Ascend

C++2.8K

runanywhere-sdks

85

Production ready toolkit to run AI locally

C++10.3K

tiny-vllm

64

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

C++1.1K

BigMoeOnEdge

77

Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp

C++548
Back to List