🌐Crawl4AI
unclecode/crawl4ai · homepage ↗
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
⭐ Hugely popular: 78k stars, gaining about 988 a week
View on GitHub ↗repo profile
momentum
durability
bus factor = how many people it takes to cover more than half the commits (6 months). 1 is a solo project; higher means the work is spread across a team. top-author share is the single busiest author's slice of those commits.
since we covered it
why it's a big deal
- Adaptive Crawling learns site patterns and stops when enough info is collected.
- Supports infinite scroll, intelligent link previews, async URL seeding, and high concurrency.
- Outputs LLM-friendly formats (Markdown, structured JSON), ideal for RAG/fine-tuning workflows.
under the hood
- Built in Python, with AsyncWebCrawler & Playwright-based browser control.
- Browser/session hooks, proxy support, stealth mode, and virtual scroll for dynamic pages.
- Modular extraction strategies: CSS/XPath, LLM-driven schema extraction, metadata, media, PDF parsingCrawl4AI democratizes AI pipelines by reimagining data harvesting as a smart, structured, and agent-friendly process - empowering developers to gather high-quality knowledge at scale.
our take from PR#13, 2025-07-23
star history
curve is sampled from GitHub's star history, plus our own daily readings since we covered it; the dashed stretch is before we first covered it, the solid line since. figures at coverage are the numbers we printed then (approx.), current count is live.
understory
Better known than its recent output, coasting a little on attention.
- output, commits & releases
- clout, star velocity
output = commits & releases; clout = star velocity, both 0 to 100 monthly indices; the gap where output runs above clout is the understory. The understory →
covered in
-
Open-Source LLM Friendly Web Crawler & Scraper
-
Open-source LLM Friendly Web Crawler & Scraper
similar projects
compare these →- 💻 BrowserUse
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
109k ACTIVE - 🕷️ Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
74k ACTIVE - 🕸️ Scrapegraph-ai
leaner, 30k stars
Python scraper based on AI
30k ACTIVE
replaces
alternatives to Apify →- Apify
market leader
closed source - Bright Data closed source
- Diffbot closed source
comments
Sign in with GitHub to add your blip on Crawl4AI.