← Back to List
⚠
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonweicj/vLLM-2080Ti-Definitive

vLLM-2080Ti-Definitive

The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 200+ tok/s single-request decode with support of FP8 weight ( Join Discord :https://discord.gg/VFqVVySdMS )

72.8/100
★ 1.0KForks: 141
View on GitHub →
Loading report...

Similar Projects

hermes-agent

90

The agent that grows with you

Python★ 249.5K

AutoGPT

96

AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

Python★ 187.6K

markitdown

88

Python tool for converting files and office documents to Markdown.

Python★ 187.3K

skills

64

Public repository for Agent Skills

Python★ 178.7K
← Back to List