Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonyoussofal/MTPLX

MTPLX

The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

84.2/100
2.4KForks: 177
View on GitHubHomepage →
Loading report...

Similar Projects

mlx-vlm

84

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

Python5.5K

vllm-metal

82

Community maintained hardware plugin for vLLM on Apple Silicon

Python1.7K

mlx-dspark

75

Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

Python673

omlx

89

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Python21.9K
Back to List