GitHub repo leaderboard by stars, growth rate and activity.
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
Scalable Python web scraping scripts for +40 popular domains
Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.
📚 This is an adapted version of Jina AI's Reader for local deployment using Docker. Convert any URL to an LLM-friendly input with a simple prefix http://127.0.0.1:3000/https://website-to-scrape.com/
Self-hosted feed generation toolkit for turning webpages, email folders, APIs, and messy web sources into reliable RSS feeds.
Lego AI Parser is an open-source application that uses OpenAI to parse visible text of HTML elements.
A Python command-line tool for scraping and downloading subtitles from AppleTV and iTunes movie pages.
Unofficial Python client for xvideos.com: search, metadata extraction, and video download (curl-cffi + selectolax)
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Self-hosted webscraper. | TypeScript | 4,905 | last pushed 11 months ago | |
| 2 | AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing. | TypeScript | 3,452 | last pushed 2 days ago | |
| 3 | Scalable Python web scraping scripts for +40 popular domains | Python | 1,079 | last pushed 1 week ago | |
| 4 | Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents. | Pascal | 843 | last pushed 2 years ago | |
| 5 | An R web crawler and scraper | R | 362 | last pushed 4 years ago | |
| 6 | 📚 This is an adapted version of Jina AI's Reader for local deployment using Docker. Convert any URL to an LLM-friendly input with a simple prefix http://127.0.0.1:3000/https://website-to-scrape.com/ | TypeScript | 307 | last pushed 1 year ago | |
| 7 | Self-hosted feed generation toolkit for turning webpages, email folders, APIs, and messy web sources into reliable RSS feeds. | TypeScript | 284 | last pushed 4 months ago | |
| 8 | Lego AI Parser is an open-source application that uses OpenAI to parse visible text of HTML elements. | Python | 237 | last pushed 2 years ago | |
| 9 | A Python command-line tool for scraping and downloading subtitles from AppleTV and iTunes movie pages. | Python | 220 | last pushed 2 days ago | |
| 10 | Unofficial Python client for xvideos.com: search, metadata extraction, and video download (curl-cffi + selectolax) | Python | 162 | last pushed 3 days ago |
All · 11,474