GitHub repo leaderboard by stars, growth rate and activity.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Declarative data automation language and Go runtime for structured extraction workflows.
The All in One Framework to Build Undefeatable Scrapers
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining"
Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change.
A simple web scraper to extract Product Data and Pricing from Amazon
Library for Rapid (Web) Crawler and Scraper Development
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25,713 | last pushed Yesterday | |
| 2 | Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. | Python | 9,500 | last pushed 2 days ago | |
| 3 | Declarative data automation language and Go runtime for structured extraction workflows. | Go | 6,009 | last pushed 24 hours ago | |
| 4 | The All in One Framework to Build Undefeatable Scrapers | Python | 5,704 | last pushed 2 months ago | |
| 5 | A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access. | JavaScript | 2,636 | last pushed 4 weeks ago | |
| 6 | Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust | Rust | 1,848 | last pushed 24 hours ago | |
| 7 | Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining" | Python | 664 | last pushed 2 years ago | |
| 8 | Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change. | Python | 482 | last pushed 2 months ago | |
| 9 | A simple web scraper to extract Product Data and Pricing from Amazon | Python | 442 | last pushed 3 years ago | |
| 10 | Library for Rapid (Web) Crawler and Scraper Development | PHP | 371 | last pushed 4 months ago |
All · 11,474