GitHub repo leaderboard by stars, growth rate and activity.
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
A Datacenter Scale Distributed Inference Serving Framework
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services | Jupyter Notebook | 18,560 | last pushed 4 months ago | |
| 2 | ✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记 | Go | 13,234 | last pushed 1 week ago | |
| 3 | LMCache: Supercharge Your LLM with the Fastest KV Cache Layer | Python | 11,733 | last pushed 23 hours ago | |
| 4 | Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. | Python | 9,562 | last pushed 2 days ago | |
| 5 | A Datacenter Scale Distributed Inference Serving Framework | Rust | 8,008 | last pushed 22 hours ago | |
| 6 | A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines | Python | 5,686 | last pushed 2 days ago | |
| 7 | A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. | Python | 5,647 | last pushed 23 hours ago | |
| 8 | Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc | Go | 5,629 | last pushed 24 hours ago | |
| 9 | Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop. | Python | 3,955 | last pushed 2 days ago | |
| 10 | High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle | Python | 3,713 | last pushed 2 weeks ago |
All · 11,474