GitHub repo leaderboard by stars, growth rate and activity.
A high-throughput and memory-efficient inference and serving engine for LLMs
The open-source AI voice studio. Clone, dictate, create.
SGLang is a high-performance serving framework for large language models and multimodal models.
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
kaldi-asr/kaldi is the official location of the Kaldi project.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
FlashInfer: Kernel Library for LLM Serving
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 91,383 | last pushed 24 hours ago | |
| 2 | The open-source AI voice studio. Clone, dictate, create. | TypeScript | 52,834 | last pushed 5 weeks ago | |
| 3 | SGLang is a high-performance serving framework for large language models and multimodal models. | Python | 35,728 | last pushed Yesterday | |
| 4 | Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability. | Rust | 15,889 | last pushed Yesterday | |
| 5 | kaldi-asr/kaldi is the official location of the Kaldi project. | Shell | 15,476 | last pushed 12 months ago | |
| 6 | CUDA on non-NVIDIA GPUs | Rust | 14,818 | last pushed 1 week ago | |
| 7 | TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. | Python | 14,580 | last pushed 24 hours ago | |
| 8 | LMCache: Supercharge Your LLM with the Fastest KV Cache Layer | Python | 11,733 | last pushed Yesterday | |
| 9 | cuDF - GPU DataFrame Library | C++ | 9,749 | last pushed 24 hours ago | |
| 10 | FlashInfer: Kernel Library for LLM Serving | Cuda | 6,363 | last pushed Yesterday |
All · 11,474