Skip to content
repository radar

// category

Search & Research

Web crawlers, scrapers, browser automation, and deep-research tools that gather and structure information at scale.

These are our picks, not a scrape of GitHub trending: every repo here is one we covered in Repository Radar because we thought it worth tracking. Ranked by stars, with live metrics and a get-it block on each.

11 repos in this category. Browse or filter the full archive → · Compare all categories →

// the space

11 repos, our 8th-largest category (Agents & Orchestration leads with 92). Combined they're up 192% since we covered them, 2nd-fastest-growing of our 9 categories. They hold 587k stars, 1k contributors and 12M weekly downloads between them. Strength runs broad, not top-heavy: the median repo (36k stars) sits close to the 53k average. 5 are breaking out and 10 still actively shipping, and Python leads on language. What's pulling ahead: crawler and scraping projects.

total

587k
stars tracked
#7 of 9
+386k
stars added since covered
#3 of 9
1k
contributors
#8 of 9
12M
weekly downloads
#3 of 9
202
commits / week
#9 of 9
5
breaking out
#8 of 9
10
still shipping
#8 of 9

per repo

36k
median stars
#1 of 9
+35k
stars added / repo
#1 of 9
+192%
growth since covered
#2 of 9
98
contributors / repo
#7 of 9
1M
downloads / repo
#2 of 9
18
commits / repo
#8 of 9
45%
% breaking out
#5 of 9
91%
% shipping
#6 of 9
🔥Firecrawl BREAKOUT ACTIVE

firecrawl/firecrawl

The API to search, scrape, and interact with the web at scale. 🔥

157k
+ 7k/mo (+32%/mo) since covered
TypeScript
product TypeScript AGPL-3.0 PR#16, PR#1
💻BrowserUse BREAKOUT ACTIVE

browser-use/browser-use

🌐 Make websites accessible for AI agents. Automate tasks online with ease.

107k
+ 4k/mo (+15%/mo) since covered
Python
library Python MIT PR#2
🌐Crawl4AI ACTIVE

unclecode/crawl4ai

🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN

75k
+ 2k/mo (+7%/mo) since covered
Python
library Python Apache-2.0 PR#13, PR#5
🕷️Scrapling BREAKOUT ACTIVE

D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

72k
+ 10k/mo (+40%/mo) since covered
Python
product Python BSD-3-Clause PR#29
🌐agent-browser BREAKOUT ACTIVE

vercel-labs/agent-browser

Browser automation CLI for AI agents

39k
+ 5k/mo (+45%/mo) since covered
Rust
product Rust Apache-2.0 PR#26
🧠Perplexica ACTIVE

ItzCrazyKns/Vane

Vane is an AI-powered answering engine.

36k
+ 1k/mo (+4%/mo) since covered
TypeScript
library TypeScript MIT PR#17
🧭Stagehand ACTIVE

browserbase/stagehand

The SDK For Browser Agents

24k
+ 818/mo (+6%/mo) since covered
TypeScript
product TypeScript MIT PR#11
🎯Skyvern ACTIVE

Skyvern-AI/skyvern

Automate browser based workflows with AI

23k
+ 605/mo (+3%/mo) since covered
Python
library Python AGPL-3.0 PR#21
🕵️DeepResearch BREAKOUT ACTIVE

Alibaba-NLP/DeepResearch

Tongyi Deep Research, the Leading Open-source Deep Research Agent

20k
+ 2k/mo (+13152%/mo) since covered
Python
Python Apache-2.0 PR#18
📚Open Deep Research FAINT SIGNAL

nickscamara/open-deep-research

An open source deep research clone. AI Agent that reasons large amounts of web data extracted with Firecrawl

6k
+ 125/mo (+3%/mo) since covered
TypeScript
TypeScript Other PR#2