GitHub repo leaderboard by stars, growth rate and activity.
A high-throughput and memory-efficient inference and serving engine for LLMs
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
A framework for efficient model inference with omni-modality models
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
Community maintained hardware plugin for vLLM on Ascend
An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.
Native Rust inference server for open models on NVIDIA GPUs. OpenAI- and Anthropic-compatible APIs, GGUF + safetensors, FP8/NVFP4/MXFP4/Q8/Q4, built-in Studio
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 91,383 | last pushed 23 hours ago | |
| 2 | The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more! | Python | 8,832 | last pushed 3 days ago | |
| 3 | A framework for efficient model inference with omni-modality models | Python | 6,743 | last pushed 23 hours ago | |
| 4 | Open-Source Personal Cloud OS for Always-On Agents | Go | 5,263 | last pushed Yesterday | |
| 5 | FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale. | Python | 4,065 | last pushed 11 months ago | |
| 6 | Community maintained hardware plugin for vLLM on Ascend | C++ | 2,786 | last pushed 23 hours ago | |
| 7 | An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact. | Go | 1,412 | last pushed 3 days ago | |
| 8 | Native Rust inference server for open models on NVIDIA GPUs. OpenAI- and Anthropic-compatible APIs, GGUF + safetensors, FP8/NVFP4/MXFP4/Q8/Q4, built-in Studio | Rust | 75 | last pushed 2 days ago |
All · 11,474