GitHub repo leaderboard by stars, growth rate and activity.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | PDF Parser for AI-ready data. Automate PDF accessibility. Open-source. | Java | 29,074 | last pushed 3 hours ago | |
| 2 | Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions. | Rust | 19,013 | last pushed 3 hours ago | |
| 3 | Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server. | Rust | 9,290 | last pushed 11 hours ago | |
| 4 | TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC | C++ | 1,063 | last pushed 3 days ago |
All · 11,658