Bảng xếp hạng repo GitHub theo sao, tốc độ tăng trưởng và mức độ hoạt động.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Scrapy, a fast high-level web crawling & scraping framework for Python.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
List of libraries, tools and APIs for web scraping and data processing.
A Chrome DevTools Protocol driver for web automation and scraping.
Distributed crawler powered by Headless Chrome
| # | Repo | Ngôn ngữ | Sao | Xu hướng 30 ngày | Cập nhật lần cuối |
|---|---|---|---|---|---|
| 1 | 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! | Python | 79.739 | push cuối 7 ngày trước | |
| 2 | Scrapy, a fast high-level web crawling & scraping framework for Python. | Python | 64.268 | push cuối 2 ngày trước | |
| 3 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25.713 | push cuối Hôm qua | |
| 4 | Elegant Scraper and Crawler Framework for Golang | Go | 25.504 | push cuối 1 tuần trước | |
| 5 | 🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥 | TypeScript | 17.403 | push cuối Hôm qua | |
| 6 | newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs: | Python | 15.151 | push cuối 1 tuần trước | |
| 7 | Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. | Python | 9.500 | push cuối 2 ngày trước | |
| 8 | List of libraries, tools and APIs for web scraping and data processing. | JavaScript | 8.146 | push cuối 3 ngày trước | |
| 9 | A Chrome DevTools Protocol driver for web automation and scraping. | Go | 7.094 | push cuối 4 tuần trước | |
| 10 | Distributed crawler powered by Headless Chrome | JavaScript | 5.635 | push cuối 3 năm trước |
Tất cả · 11.474