GitHub repo leaderboard by stars, growth rate and activity.
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Easy to use lightweight web crawler(易用的轻量化网络爬虫)
A Kotlin-based testing/scraping/parsing library providing the ability to analyze and extract data from HTML (server & client-side rendered). It places particular emphasis on ease of use and a high level of readability by providing an intuitive DSL. It aims to be a testing lib, but can also be used to scrape websites in a convenient fashion.
Headless/full Java browser with support for downloading files, working with cookies, retrieving HTML and simulating real user input. Possible via Node.js with Puppeteer and/or Playwright. Main focus on ease of use and high-level methods.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety. | Java | 11,389 | last pushed Yesterday | |
| 2 | 新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。 | Java | 11,351 | last pushed 3 years ago | |
| 3 | Easy to use lightweight web crawler(易用的轻量化网络爬虫) | Java | 2,511 | last pushed 8 months ago | |
| 4 | A Kotlin-based testing/scraping/parsing library providing the ability to analyze and extract data from HTML (server & client-side rendered). It places particular emphasis on ease of use and a high level of readability by providing an intuitive DSL. It aims to be a testing lib, but can also be used to scrape websites in a convenient fashion. | Kotlin | 875 | last pushed 4 months ago | |
| 5 | Android 本地网络小说爬虫,基于jsoup及xpath | Java | 403 | last pushed 6 years ago | |
| 6 | 本仓库收集整理爬虫相关资源,开发语言以Java为主 | — | 323 | last pushed 6 years ago | |
| 7 | Headless/full Java browser with support for downloading files, working with cookies, retrieving HTML and simulating real user input. Possible via Node.js with Puppeteer and/or Playwright. Main focus on ease of use and high-level methods. | Java | 127 | last pushed 2 years ago |
All · 11,474