GitHub repo leaderboard by stars, growth rate and activity.
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
A Datacenter Scale Distributed Inference Serving Framework
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A programmable Mixture-of-Models router for heterogeneous LLM inference
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services | Jupyter Notebook | 18,560 | last pushed 4 months ago | |
| 2 | ✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记 | Go | 13,234 | last pushed 1 week ago | |
| 3 | LMCache: Supercharge Your LLM with the Fastest KV Cache Layer | Python | 11,749 | last pushed 49 minutes ago | |
| 4 | Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. | Python | 9,564 | last pushed 1 hour ago | |
| 5 | A Datacenter Scale Distributed Inference Serving Framework | Rust | 8,021 | last pushed 32 minutes ago | |
| 6 | Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. | C++ | 6,552 | — | last pushed 33 minutes ago |
| 7 | A programmable Mixture-of-Models router for heterogeneous LLM inference | Go | 5,722 | — | last pushed 43 minutes ago |
| 8 | A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines | Python | 5,688 | last pushed 2 days ago | |
| 9 | A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. | Python | 5,664 | last pushed 19 hours ago | |
| 10 | Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc | Go | 5,636 | last pushed 21 hours ago |
All · 11,658