GitHub repo leaderboard by stars, growth rate and activity.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
A fast, helpful, and open-source document parser
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages. | Python | 89,222 | last pushed 2 months ago | |
| 2 | Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows. | Python | 79,590 | last pushed 2 days ago | |
| 3 | PDF Parser for AI-ready data. Automate PDF accessibility. Open-source. | Java | 29,055 | last pushed 2 days ago | |
| 4 | Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions. | Rust | 18,983 | last pushed Yesterday | |
| 5 | A fast, helpful, and open-source document parser | Rust | 12,278 | last pushed Yesterday | |
| 6 | Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages. | Rust | 1,773 | last pushed 2 years ago |
All · 11,474