GitHub repo leaderboard by stars, growth rate and activity.
The context API to search, scrape, and interact with the web at scale. 🔥
Create agents that monitor and act on your behalf. Your agents are standing by!
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
👾 Fast and simple video download library and CLI tool written in Go
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | The context API to search, scrape, and interact with the web at scale. 🔥 | TypeScript | 178,457 | last pushed Yesterday | |
| 2 | Create agents that monitor and act on your behalf. Your agents are standing by! | Ruby | 49,920 | last pushed 2 days ago | |
| 3 | A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。 | JavaScript | 44,521 | last pushed Yesterday | |
| 4 | 👾 Fast and simple video download library and CLI tool written in Go | Go | 31,670 | last pushed 6 months ago | |
| 5 | The fast, flexible, and elegant library for parsing and manipulating HTML and XML. | TypeScript | 30,482 | last pushed 23 hours ago | |
| 6 | Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. | TypeScript | 25,713 | last pushed Yesterday | |
| 7 | Elegant Scraper and Crawler Framework for Golang | Go | 25,504 | last pushed 1 week ago | |
| 8 | 🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。 | Python | 20,035 | last pushed 11 months ago | |
| 9 | 🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥 | TypeScript | 17,403 | last pushed Yesterday | |
| 10 | newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs: | Python | 15,151 | last pushed 1 week ago |
All · 11,474