GitHub repo leaderboard by stars, growth rate and activity.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Faster Whisper transcription with CTranslate2
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
A vector index built on TurboQuant, written in Rust with Python bindings
Accessible large language models via k-bit quantization for PyTorch.
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Your Cheat Sheet for AI Engineering Interview – Questions and Answers.
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024) | Python | 74,682 | last pushed 2 days ago | |
| 2 | Faster Whisper transcription with CTranslate2 | Python | 25,310 | last pushed 10 months ago | |
| 3 | 中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs) | Python | 18,942 | last pushed 5 months ago | |
| 4 | A vector index built on TurboQuant, written in Rust with Python bindings | Rust | 16,734 | last pushed 3 weeks ago | |
| 5 | Accessible large language models via k-bit quantization for PyTorch. | Python | 8,466 | last pushed 3 days ago | |
| 6 | A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU. | C | 7,382 | last pushed 2 weeks ago | |
| 7 | Your Cheat Sheet for AI Engineering Interview – Questions and Answers. | Markdown | 3,037 | last pushed Yesterday | |
| 8 | Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks | Python | 1,238 | last pushed 3 days ago |
All · 11,474