GitHub repo leaderboard by stars, growth rate and activity.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
A community-supported supercharged document management system: scan, index and archive all your documents
ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images, text, and various file types to a wide range of destinations.
Pure Javascript OCR for more than 100 Languages 📖🎉🖥
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages. | Python | 89,222 | last pushed 2 months ago | |
| 2 | Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows. | Python | 79,590 | last pushed 2 days ago | |
| 3 | OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。 | Python | 47,223 | last pushed 10 months ago | |
| 4 | A community-supported supercharged document management system: scan, index and archive all your documents | Python | 44,967 | last pushed 24 hours ago | |
| 5 | ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images, text, and various file types to a wide range of destinations. | C# | 39,508 | last pushed Yesterday | |
| 6 | Pure Javascript OCR for more than 100 Languages 📖🎉🖥 | JavaScript | 38,695 | last pushed 4 months ago | |
| 7 | OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched | Python | 34,705 | last pushed 2 days ago | |
| 8 | Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc. | Python | 29,988 | last pushed 9 months ago | |
| 9 | PDF Parser for AI-ready data. Automate PDF accessibility. Open-source. | Java | 29,055 | last pushed 2 days ago | |
| 10 | 🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition. | JavaScript | 19,437 | last pushed 2 months ago |
All · 11,474