GitHub repo leaderboard by stars, growth rate and activity.
A high-throughput and memory-efficient inference and serving engine for LLMs
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Superduper: End-to-end framework for building custom AI applications and agents.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 91,451 | last pushed 2 hours ago | |
| 2 | Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads. | Python | 43,770 | last pushed 4 hours ago | |
| 3 | 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地) | HTML | 25,031 | last pushed 2 months ago | |
| 4 | TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. | Python | 14,592 | last pushed 2 hours ago | |
| 5 | Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud. | Python | 12,528 | last pushed 4 days ago | |
| 6 | The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster. | Python | 10,585 | last pushed 2 hours ago | |
| 7 | The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more! | Python | 8,834 | last pushed 3 days ago | |
| 8 | A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. | Python | 5,664 | last pushed 20 hours ago | |
| 9 | Superduper: End-to-end framework for building custom AI applications and agents. | Python | 5,320 | last pushed 1 year ago | |
| 10 | High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle | Python | 3,714 | last pushed 2 weeks ago |
All · 11,658