GitHub repo leaderboard by stars, growth rate and activity.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
TokenSpeed is a speed-of-light LLM inference engine.
Native Rust inference server for open models on NVIDIA GPUs. OpenAI- and Anthropic-compatible APIs, GGUF + safetensors, FP8/NVFP4/MXFP4/Q8/Q4, built-in Studio
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 91,383 | last pushed 22 hours ago | |
| 2 | SGLang is a high-performance serving framework for large language models and multimodal models. | Python | 35,728 | last pushed 23 hours ago | |
| 3 | TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. | Python | 14,580 | last pushed 22 hours ago | |
| 4 | TokenSpeed is a speed-of-light LLM inference engine. | Python | 2,112 | last pushed 22 hours ago | |
| 5 | Native Rust inference server for open models on NVIDIA GPUs. OpenAI- and Anthropic-compatible APIs, GGUF + safetensors, FP8/NVFP4/MXFP4/Q8/Q4, built-in Studio | Rust | 75 | last pushed 2 days ago |
All · 11,474