GitHub repo leaderboard by stars, growth rate and activity.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Web Crawler/Spider for NodeJS + server-side jQuery ;-)
Crawly, a high-level web crawling & scraping framework for Elixir.
Extract structured data from web sites. Web sites scraping.
Turn Webpage to LLM friendly input text. Similar to Firecrawl and Jina Reader API. Makes RAG, AI web scraping, image & webpage links extraction easy.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows. | Python | 79,590 | last pushed 2 days ago | |
| 2 | Web Crawler/Spider for NodeJS + server-side jQuery ;-) | TypeScript | 6,793 | last pushed 3 months ago | |
| 3 | Crawly, a high-level web crawling & scraping framework for Elixir. | Elixir | 1,116 | last pushed 1 year ago | |
| 4 | Extract structured data from web sites. Web sites scraping. | Go | 716 | last pushed 4 years ago | |
| 5 | Turn Webpage to LLM friendly input text. Similar to Firecrawl and Jina Reader API. Makes RAG, AI web scraping, image & webpage links extraction easy. | Python | 305 | last pushed 7 months ago |
All · 11,658