Skip to content
repository radar
Search & Research
Repo of the Month Apr 2025 · +16%

🌐Crawl4AI

ACTIVE STEADY

unclecode/crawl4ai · homepage ↗

🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN

⭐ Hugely popular: 78k stars, gaining about 988 a week

View on GitHub ↗

repo profile

vintage 2024 2 years old
delivery library
language Python
license Apache-2.0

momentum

total stars 78k
stars added last week +988
weekly downloads 390k /wk · PyPI
used by 3k repos & packages
commits / week 0 cooling · -100% vs prior mo
issues closed 96% ~24d to close, median
contributors 91
release cadence monthly
last activity today

durability

backing community / independent
openness permissive
bus factor 2 concentrated
top-author share 45% 6 mo

bus factor = how many people it takes to cover more than half the commits (6 months). 1 is a solo project; higher means the work is spread across a team. top-author share is the single busiest author's slice of those commits.

since we covered it

monthly average + 3k/mo (+7%/mo) · + 43k total since PR#5

why it's a big deal

  • Adaptive Crawling learns site patterns and stops when enough info is collected.
  • Supports infinite scroll, intelligent link previews, async URL seeding, and high concurrency.
  • Outputs LLM-friendly formats (Markdown, structured JSON), ideal for RAG/fine-tuning workflows.

under the hood

  • Built in Python, with AsyncWebCrawler & Playwright-based browser control.
  • Browser/session hooks, proxy support, stealth mode, and virtual scroll for dynamic pages.
  • Modular extraction strategies: CSS/XPath, LLM-driven schema extraction, metadata, media, PDF parsingCrawl4AI democratizes AI pipelines by reimagining data harvesting as a smart, structured, and agent-friendly process - empowering developers to gather high-quality knowledge at scale.

our take from PR#13, 2025-07-23

star history

PR#5 · 36k PR#13 · 48k78k now May 2024Aug 2026
  1. PR#5 36k 2025-04-02
  2. PR#13 48k 2025-07-23
  3. now 78k + 43k since first covered

curve is sampled from GitHub's star history, plus our own daily readings since we covered it; the dashed stretch is before we first covered it, the solid line since. figures at coverage are the numbers we printed then (approx.), current count is live.

understory

Better known than its recent output, coasting a little on attention.

-11 understory score output 69 · clout 80
Aug 2025 Jul 2026
  • output, commits & releases
  • clout, star velocity

output = commits & releases; clout = star velocity, both 0 to 100 monthly indices; the gap where output runs above clout is the understory. The understory →

covered in

  • PR#13 2025-07-23 above the radar

    Open-Source LLM Friendly Web Crawler & Scraper

  • PR#5 2025-04-02 on the radar

    Open-source LLM Friendly Web Crawler & Scraper

similar projects

compare these →
  • 💻 BrowserUse

    🌐 Make websites accessible for AI agents. Automate tasks online with ease.

    109k ACTIVE
  • 🕷️ Scrapling

    🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

    74k ACTIVE
  • 🕸️ Scrapegraph-ai

    leaner, 30k stars

    Python scraper based on AI

    30k ACTIVE
  • Apify

    market leader

    closed source
  • Bright Data
    closed source
  • Diffbot
    closed source

comments

Sign in with GitHub to add your blip on Crawl4AI.

loading comments…