GitHub repo leaderboard by stars, growth rate and activity.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Scrapy, a fast high-level web crawling & scraping framework for Python.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
List of libraries, tools and APIs for web scraping and data processing.
A Chrome DevTools Protocol driver for web automation and scraping.
Distributed crawler powered by Headless Chrome
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! | Python | 80,073 | last pushed 7 days ago | |
| 2 | Scrapy, a fast high-level web crawling & scraping framework for Python. | Python | 64,282 | last pushed 13 hours ago | |
| 3 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25,731 | last pushed 16 hours ago | |
| 4 | Elegant Scraper and Crawler Framework for Golang | Go | 25,507 | last pushed 1 week ago | |
| 5 | 🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥 | TypeScript | 17,415 | last pushed 4 hours ago | |
| 6 | newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs: | Python | 15,153 | last pushed 1 week ago | |
| 7 | Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. | Python | 9,508 | last pushed 22 hours ago | |
| 8 | List of libraries, tools and APIs for web scraping and data processing. | JavaScript | 8,147 | last pushed 3 days ago | |
| 9 | A Chrome DevTools Protocol driver for web automation and scraping. | Go | 7,096 | last pushed 4 weeks ago | |
| 10 | Distributed crawler powered by Headless Chrome | JavaScript | 5,635 | last pushed 3 years ago |
All · 11,658