GitHub repo leaderboard by stars, growth rate and activity.
The context API to search, scrape, and interact with the web at scale. 🔥
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
A collection of awesome web crawler,spider in different languages
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
The Ultimate Information Gathering Toolkit
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | The context API to search, scrape, and interact with the web at scale. 🔥 | TypeScript | 178,074 | — | last pushed 2 days ago |
| 2 | Python scraper based on AI | Python | 30,749 | — | last pushed 3 days ago |
| 3 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25,701 | — | last pushed 2 days ago |
| 4 | Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架 | Go | 12,269 | — | last pushed 7 months ago |
| 5 | 新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。 | Java | 11,352 | — | last pushed 3 years ago |
| 6 | Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. | Python | 9,493 | — | last pushed 3 days ago |
| 7 | 🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients. | JavaScript | 7,422 | — | last pushed 3 days ago |
| 8 | A collection of awesome web crawler,spider in different languages | — | 7,304 | — | last pushed 2 years ago |
| 9 | The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta. | TypeScript | 5,196 | — | last pushed 2 days ago |
| 10 | The Ultimate Information Gathering Toolkit | Python | 4,120 | — | last pushed 9 months ago |
All · 11,474