GitHub repo leaderboard by stars, growth rate and activity.
INFO-SPIDER 是一个集众多数据源于一身的爬虫工具箱🧰,旨在安全快捷的帮助用户拿回自己的数据,工具代码开源,流程透明。支持数据源包括GitHub、QQ邮箱、网易邮箱、阿里邮箱、新浪邮箱、Hotmail邮箱、Outlook邮箱、京东、淘宝、支付宝、中国移动、中国联通、中国电信、知乎、哔哩哔哩、网易云音乐、QQ好友、QQ群、生成朋友圈相册、浏览器浏览历史、12306、博客园、CSDN博客、开源中国博客、简书。
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
Python爬虫实战 - 模拟登陆各大网站 包含但不限于:滑块验证、拼多多、美团、百度、bilibili、大众点评、淘宝,如果喜欢请start ❤️
Flexible Node.js AI-assisted crawler library
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Advanced python library to scrap Twitter (tweets, users) from unofficial API
🕵️ Python project to crawl for JavaScript files and search for secrets like API keys, authorization tokens, hardcoded credentials, etc.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | INFO-SPIDER 是一个集众多数据源于一身的爬虫工具箱🧰,旨在安全快捷的帮助用户拿回自己的数据,工具代码开源,流程透明。支持数据源包括GitHub、QQ邮箱、网易邮箱、阿里邮箱、新浪邮箱、Hotmail邮箱、Outlook邮箱、京东、淘宝、支付宝、中国移动、中国联通、中国电信、知乎、哔哩哔哩、网易云音乐、QQ好友、QQ群、生成朋友圈相册、浏览器浏览历史、12306、博客园、CSDN博客、开源中国博客、简书。 | Python | 8,249 | last pushed 5 months ago | |
| 2 | AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing. | TypeScript | 3,453 | last pushed 2 days ago | |
| 3 | Python爬虫实战 - 模拟登陆各大网站 包含但不限于:滑块验证、拼多多、美团、百度、bilibili、大众点评、淘宝,如果喜欢请start ❤️ | Python | 3,386 | last pushed 3 years ago | |
| 4 | Flexible Node.js AI-assisted crawler library | TypeScript | 1,879 | last pushed 7 hours ago | |
| 5 | The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns | Python | 1,611 | last pushed 1 year ago | |
| 6 | 浏览过的精彩逆向文章汇总,值得一看 | — | 1,414 | last pushed 5 months ago | |
| 7 | A Facebook crawler | Python | 694 | last pushed 6 years ago | |
| 8 | Advanced python library to scrap Twitter (tweets, users) from unofficial API | Python | 621 | last pushed 3 years ago | |
| 9 | 🕵️ Python project to crawl for JavaScript files and search for secrets like API keys, authorization tokens, hardcoded credentials, etc. | Python | 437 | last pushed 5 months ago | |
| 10 | Crawl telegra.ph searching for nudes! | Python | 376 | last pushed 2 years ago |
All · 11,658