Tools for merging pretrained large language models.
A high-throughput and memory-efficient inference and serving engine for LLMs
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
SGLang is a high-performance serving framework for large language models and multimodal models.