GitHub repo leaderboard by stars, growth rate and activity.
A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
Maid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely.
Atomic Agent is a local-first AI agent. Runs open-weight models on your own machine via llama.cpp.
A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
Self-hosted AI workspace where chat becomes visual workflows, multi-agent operations, and reviewable automations. Local memory; local or cloud models
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support! | TypeScript | 10,938 | last pushed 2 years ago | |
| 2 | Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. | Python | 9,562 | last pushed 2 days ago | |
| 3 | Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. | Python | 6,346 | last pushed 23 hours ago | |
| 4 | The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device. | TypeScript | 3,074 | last pushed 23 hours ago | |
| 5 | Maid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely. | TypeScript | 2,670 | last pushed 3 days ago | |
| 6 | Atomic Agent is a local-first AI agent. Runs open-weight models on your own machine via llama.cpp. | TypeScript | 2,516 | last pushed Yesterday | |
| 7 | A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows | Python | 2,243 | last pushed 1 week ago | |
| 8 | Create characters in Unity with LLMs! | C# | 1,707 | last pushed 4 months ago | |
| 9 | Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG | C++ | 1,543 | last pushed Yesterday | |
| 10 | Self-hosted AI workspace where chat becomes visual workflows, multi-agent operations, and reviewable automations. Local memory; local or cloud models | TypeScript | 100 | last pushed 1 week ago |
All · 11,474