Bảng xếp hạng repo GitHub theo sao, tốc độ tăng trưởng và mức độ hoạt động.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
TokenSpeed is a speed-of-light LLM inference engine.
Native Rust inference server for open models on NVIDIA GPUs. OpenAI- and Anthropic-compatible APIs, GGUF + safetensors, FP8/NVFP4/MXFP4/Q8/Q4, built-in Studio
| # | Repo | Ngôn ngữ | Sao | Xu hướng 30 ngày | Cập nhật lần cuối |
|---|---|---|---|---|---|
| 1 | A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 91.383 | push cuối 23 giờ trước | |
| 2 | SGLang is a high-performance serving framework for large language models and multimodal models. | Python | 35.728 | push cuối 24 giờ trước | |
| 3 | TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. | Python | 14.580 | push cuối 23 giờ trước | |
| 4 | TokenSpeed is a speed-of-light LLM inference engine. | Python | 2.112 | push cuối 23 giờ trước | |
| 5 | Native Rust inference server for open models on NVIDIA GPUs. OpenAI- and Anthropic-compatible APIs, GGUF + safetensors, FP8/NVFP4/MXFP4/Q8/Q4, built-in Studio | Rust | 75 | push cuối 2 ngày trước |
Tất cả · 11.474