Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
C++jmaczan/tiny-vllm

tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

63.5/100
1.1KForks: 88
View on GitHub
Loading report...

Similar Projects

uccl

79

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

C++1.5K

cactus

83

Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.

C++6.0K

ZhiLight

53

A highly optimized LLM inference acceleration engine for Llama and its variants.

C++908

LLM-Hub

79

Local LLM, image&video&music generator, vibecode like cursor with local models on your phone

C++577
Back to List