GitHub repo leaderboard by stars, growth rate and activity.
FlashInfer: Kernel Library for LLM Serving
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Achieve state of the art inference performance with modern accelerators on Kubernetes
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | FlashInfer: Kernel Library for LLM Serving | Cuda | 6,372 | last pushed 41 minutes ago | |
| 2 | A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. | Python | 5,664 | last pushed 19 hours ago | |
| 3 | Achieve state of the art inference performance with modern accelerators on Kubernetes | Shell | 4,495 | last pushed 20 minutes ago |
All · 11,658