🎙️FunASR
modelscope/FunASR · homepage ↗
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
⭐ Very popular: 20k stars, gaining about 144 a week
View on GitHub ↗repo profile
momentum
durability
bus factor = how many people it takes to cover more than half the commits (6 months). 1 is a solo project; higher means the work is spread across a team. top-author share is the single busiest author's slice of those commits.
since we covered it
why it's a big deal
- Industrial-Grade Models: FunASR provides access to models like Paraformer, trained on extensive datasets (e.g., 60,000 hours of Mandarin speech), delivering high accuracy and efficiency. .
- Comprehensive Feature Set: Beyond ASR, it includes VAD, punctuation restoration, speaker verification, and diarization, enabling robust speech processing pipelines. .
- Flexible Deployment: Supports both offline and real-time transcription services, with deployment options via Docker, making it adaptable to various environments. .
under the hood
- Multi-Platform Support: Compatible with CPU and GPU inference through ONNX, libtorch, and TensorRT, catering to diverse hardware setups. .
- Extensive Model Zoo: Offers a wide range of pretrained models accessible via ModelScope and Hugging Face, covering multiple languages and tasks. .
- Developer-Friendly: Provides comprehensive documentation, tutorials, and scripts for training, fine-tuning, and deploying models, facilitating ease of use. .
our take from PR#7, 2025-04-30
star history
- PR#7 10k 2025-04-30
- now 20k + 10k since first covered
curve is sampled from GitHub's star history, plus our own daily readings since we covered it; the dashed stretch is before we first covered it, the solid line since. figures at coverage are the numbers we printed then (approx.), current count is live.
understory
Quietly building: more output than attention, for now.
- output, commits & releases
- clout, star velocity
output = commits & releases; clout = star velocity, both 0 to 100 monthly indices; the gap where output runs above clout is the understory. The understory →
covered in
-
A Comprehensive End-to-End Speech Recognition Toolkit
similar projects
compare these →- 🎙️ meetily
Rust
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes
29k ACTIVE - 🎙️ Real-Time-Voice-Cloning
3× the stars
Clone a voice in 5 seconds to generate arbitrary speech in real-time
60k ACTIVE - 👁️ supervision
2.5× the stars
We write your reusable computer vision tools. 💜
49k ACTIVE
comments
Sign in with GitHub to add your blip on FunASR.