Bảng xếp hạng repo GitHub theo sao, tốc độ tăng trưởng và mức độ hoạt động.
The context API to search, scrape, and interact with the web at scale. 🔥
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
A collection of awesome web crawler,spider in different languages
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
| # | Repo | Ngôn ngữ | Sao | Xu hướng 30 ngày | Cập nhật lần cuối |
|---|---|---|---|---|---|
| 1 | The context API to search, scrape, and interact with the web at scale. 🔥 | TypeScript | 178.457 | push cuối Hôm qua | |
| 2 | Python scraper based on AI | Python | 30.783 | push cuối 4 ngày trước | |
| 3 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25.713 | push cuối Hôm qua | |
| 4 | Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架 | Go | 12.270 | push cuối 7 tháng trước | |
| 5 | 新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。 | Java | 11.351 | push cuối 3 năm trước | |
| 6 | Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. | Python | 9.500 | push cuối 2 ngày trước | |
| 7 | 🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients. | JavaScript | 7.427 | push cuối Hôm qua | |
| 8 | A collection of awesome web crawler,spider in different languages | — | 7.304 | push cuối 2 năm trước | |
| 9 | The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta. | TypeScript | 5.211 | push cuối 2 ngày trước | |
| 10 | The Ultimate Information Gathering Toolkit | Python | 4.123 | push cuối 9 tháng trước |
Tất cả · 11.474