GitHub repo leaderboard by stars, growth rate and activity.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Agent for collecting, processing, aggregating, and writing metrics, logs, and other arbitrary data.
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.
豆瓣电影top250、斗鱼爬取json数据以及爬取美女图片、淘宝、有缘、CrawlSpider爬取红娘网相亲人的部分基本信息以及红娘网分布式爬取和存储redis、爬虫小demo、Selenium、爬取多点、django开发接口、爬取有缘网信息、模拟知乎登录、模拟github登录、模拟图虫网登录、爬取多点商城整站数据、爬取微信公众号历史文章、爬取微信群或者微信好友分享的文章、itchat监听指定微信公众号分享的文章
Undetected web-scraping & seamless HTML parsing in Python!
dude uncomplicated data extraction: A simple framework for writing web scrapers using Python decorators
High-performance HTML5 parser for Ruby based on Lexbor, with support for both CSS selectors and XPath.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! | Python | 79,739 | last pushed 7 days ago | |
| 2 | Agent for collecting, processing, aggregating, and writing metrics, logs, and other arbitrary data. | Go | 17,798 | last pushed Yesterday | |
| 3 | jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety. | Java | 11,389 | last pushed Yesterday | |
| 4 | 新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。 | Java | 11,351 | last pushed 3 years ago | |
| 5 | 基于appium的app自动遍历工具 | Scala | 1,237 | last pushed 5 months ago | |
| 6 | Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents. | Pascal | 843 | last pushed 2 years ago | |
| 7 | 豆瓣电影top250、斗鱼爬取json数据以及爬取美女图片、淘宝、有缘、CrawlSpider爬取红娘网相亲人的部分基本信息以及红娘网分布式爬取和存储redis、爬虫小demo、Selenium、爬取多点、django开发接口、爬取有缘网信息、模拟知乎登录、模拟github登录、模拟图虫网登录、爬取多点商城整站数据、爬取微信公众号历史文章、爬取微信群或者微信好友分享的文章、itchat监听指定微信公众号分享的文章 | Python | 777 | last pushed 4 years ago | |
| 8 | Undetected web-scraping & seamless HTML parsing in Python! | Python | 563 | last pushed 5 months ago | |
| 9 | dude uncomplicated data extraction: A simple framework for writing web scrapers using Python decorators | Python | 426 | last pushed 1 year ago | |
| 10 | High-performance HTML5 parser for Ruby based on Lexbor, with support for both CSS selectors and XPath. | C | 414 | last pushed 2 months ago |
All · 11,474