GitHub repo leaderboard by stars, growth rate and activity.
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
A Datacenter Scale Distributed Inference Serving Framework
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Official SGLang × Datawhale course on LLM inference (中英双语): understand inference, build a mini-sglang from scratch, then read the real SGLang source and land your first PR. 《从零手搓SGLang》:读懂推理,手搓 mini-sglang,吃透 SGLang 源码。
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. | Python | 9,562 | last pushed 2 days ago | |
| 2 | A Datacenter Scale Distributed Inference Serving Framework | Rust | 8,008 | last pushed 22 hours ago | |
| 3 | A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. | Python | 5,647 | last pushed 22 hours ago | |
| 4 | Official SGLang × Datawhale course on LLM inference (中英双语): understand inference, build a mini-sglang from scratch, then read the real SGLang source and land your first PR. 《从零手搓SGLang》:读懂推理,手搓 mini-sglang,吃透 SGLang 源码。 | Python | 637 | last pushed 2 days ago |
All · 11,474