Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
HTMLpatchy631/time-to-first-token

time-to-first-token

A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

67.9/100
618Forks: 77
View on GitHub
Loading report...

Similar Projects

llm-action

73

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

HTML24.9K

prompts.chat

85

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.

HTML167.2K

llm_interview_note

56

主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题

HTML14.9K

AgentGuide

61

https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG | 转行大模型 | 大模型面试 | 算法工程师 | 面试题库 | 强化学习|数据合成

HTML8.3K
Back to List