Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
PythonFlashML-org/FreeToken

FreeToken

FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

79.6/100
12.0KForks: 1.2K
View on GitHubHomepage →
Loading report...

Similar Projects

sglang

91

SGLang is a high-performance serving framework for large language models and multimodal models.

Python35.6K

vllm

93

A high-throughput and memory-efficient inference and serving engine for LLMs

Python91.2K

nano-vllm

51

Nano vLLM

Python15.3K

text-generation-inference

68

Large Language Model Text Generation Inference

Python10.9K
Back to List