← Back to List
⚠
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
HTMLpatchy631/time-to-first-token

time-to-first-token

A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

53.3/100
★ 964Forks: 108
View on GitHub →
Loading report...

Similar Projects

llm-action

66

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

HTML★ 25.1K

prompts.chat

79

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.

HTML★ 171.6K

llm_interview_note

52

主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题

HTML★ 15.2K

awesome-openclaw-agents

79

162 production-ready AI agent templates for OpenClaw. SOUL.md configs across 19 categories. Submit yours!

HTML★ 4.0K
← Back to List