Bảng xếp hạng repo GitHub theo sao, tốc độ tăng trưởng và mức độ hoạt động.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Declarative data automation language and Go runtime for structured extraction workflows.
The All in One Framework to Build Undefeatable Scrapers
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining"
Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change.
A simple web scraper to extract Product Data and Pricing from Amazon
Library for Rapid (Web) Crawler and Scraper Development
| # | Repo | Ngôn ngữ | Sao | Xu hướng 30 ngày | Cập nhật lần cuối |
|---|---|---|---|---|---|
| 1 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25.731 | push cuối 16 giờ trước | |
| 2 | Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. | Python | 9.508 | push cuối 23 giờ trước | |
| 3 | Declarative data automation language and Go runtime for structured extraction workflows. | Go | 6.009 | push cuối 3 giờ trước | |
| 4 | The All in One Framework to Build Undefeatable Scrapers | Python | 5.706 | push cuối 2 tháng trước | |
| 5 | A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access. | JavaScript | 2.635 | push cuối 4 tuần trước | |
| 6 | Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust | Rust | 1.891 | push cuối 1 giờ trước | |
| 7 | Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining" | Python | 664 | push cuối 2 năm trước | |
| 8 | Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change. | Python | 485 | push cuối 2 tháng trước | |
| 9 | A simple web scraper to extract Product Data and Pricing from Amazon | Python | 442 | push cuối 3 năm trước | |
| 10 | Library for Rapid (Web) Crawler and Scraper Development | PHP | 371 | push cuối 4 tháng trước |
Tất cả · 11.658