GitHub repo leaderboard by stars, growth rate and activity.
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
Web Crawler/Spider for NodeJS + server-side jQuery ;-)
Scrape metadata from any URL using Open Graph, JSON-LD, HTML meta tags, and smart fallbacks.
web spider built by puppeteer, support task-queue and task-scheduling by decorators,support nedb / mongodb, support data visualization; 基于puppeteer的web爬虫框架,提供灵活的任务队列管理调度方案,提供便捷的数据保存方案(nedb/mongodb),提供数据可视化和用户交互的实现方案
An MCP server providing web search and content extraction capabilities. Integrates DuckDuckGo search functionality and URL content extraction into your MCP environment, enabling AI assistants to search the web and extract webpage content.
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | The fast, flexible, and elegant library for parsing and manipulating HTML and XML. | TypeScript | 30,482 | last pushed 23 hours ago | |
| 2 | Web Crawler/Spider for NodeJS + server-side jQuery ;-) | TypeScript | 6,793 | last pushed 3 months ago | |
| 3 | Scrape metadata from any URL using Open Graph, JSON-LD, HTML meta tags, and smart fallbacks. | HTML | 2,736 | last pushed 3 weeks ago | |
| 4 | web spider built by puppeteer, support task-queue and task-scheduling by decorators,support nedb / mongodb, support data visualization; 基于puppeteer的web爬虫框架,提供灵活的任务队列管理调度方案,提供便捷的数据保存方案(nedb/mongodb),提供数据可视化和用户交互的实现方案 | TypeScript | 338 | last pushed 5 years ago | |
| 5 | An MCP server providing web search and content extraction capabilities. Integrates DuckDuckGo search functionality and URL content extraction into your MCP environment, enabling AI assistants to search the web and extract webpage content. | JavaScript | 133 | last pushed 2 months ago |
All · 11,474