GitHub repo leaderboard by stars, growth rate and activity.
A high-throughput and memory-efficient inference and serving engine for LLMs
🔥 MaxKB is an open-source platform for building enterprise-grade agents. 强大易用的开源企业级智能体平台。
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.
A native macOS app that allows users to chat with a local LLM that can respond with information from files, folders and websites on your Mac without installing any other software. Powered by llama.cpp.
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | A high-throughput and memory-efficient inference and serving engine for LLMs | Python | 91,383 | last pushed 23 hours ago | |
| 2 | 🔥 MaxKB is an open-source platform for building enterprise-grade agents. 强大易用的开源企业级智能体平台。 | Python | 22,750 | last pushed 23 hours ago | |
| 3 | Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025). | Python | 15,567 | last pushed 2 days ago | |
| 4 | Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code | Rust | 8,358 | last pushed Yesterday | |
| 5 | Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python. | Python | 8,259 | last pushed 10 months ago | |
| 6 | A native macOS app that allows users to chat with a local LLM that can respond with information from files, folders and websites on your Mac without installing any other software. Powered by llama.cpp. | Swift | 3,311 | last pushed 4 months ago | |
| 7 | Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG | C++ | 1,543 | last pushed Yesterday | |
| 8 | Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks | Python | 1,238 | last pushed 3 days ago |
All · 11,474