Skip to content
repository radar
Voice, Vision & Multimodal

🎙️FunASR

ACTIVE STEADY

modelscope/FunASR · homepage ↗

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

⭐ Very popular: 20k stars, gaining about 144 a week

View on GitHub ↗

repo profile

vintage 2022 4 years old
delivery product
language Python
license MIT

momentum

total stars 20k
stars added last week +144
binary downloads 252 total · GitHub releases
used by 1k repos & packages
commits / week 53 steady · -1% vs prior mo
issues closed 100% ~19d to close, median
contributors 194 (+2 in 15d)
release cadence weekly
last activity yesterday

durability

backing company-owned Alibaba
openness permissive
bus factor 1 solo
top-author share 94% 6 mo

bus factor = how many people it takes to cover more than half the commits (6 months). 1 is a solo project; higher means the work is spread across a team. top-author share is the single busiest author's slice of those commits.

since we covered it

monthly average + 619/mo (+6%/mo) · + 10k total since PR#7

why it's a big deal

  • Industrial-Grade Models: FunASR provides access to models like Paraformer, trained on extensive datasets (e.g., 60,000 hours of Mandarin speech), delivering high accuracy and efficiency. ​.
  • Comprehensive Feature Set: Beyond ASR, it includes VAD, punctuation restoration, speaker verification, and diarization, enabling robust speech processing pipelines. ​.
  • Flexible Deployment: Supports both offline and real-time transcription services, with deployment options via Docker, making it adaptable to various environments. ​.

under the hood

  • Multi-Platform Support: Compatible with CPU and GPU inference through ONNX, libtorch, and TensorRT, catering to diverse hardware setups. ​.
  • Extensive Model Zoo: Offers a wide range of pretrained models accessible via ModelScope and Hugging Face, covering multiple languages and tasks. ​.
  • Developer-Friendly: Provides comprehensive documentation, tutorials, and scripts for training, fine-tuning, and deploying models, facilitating ease of use. ​.

our take from PR#7, 2025-04-30

star history

PR#7 · 10k20k now Nov 2022Aug 2026
  1. PR#7 10k 2025-04-30
  2. now 20k + 10k since first covered

curve is sampled from GitHub's star history, plus our own daily readings since we covered it; the dashed stretch is before we first covered it, the solid line since. figures at coverage are the numbers we printed then (approx.), current count is live.

understory

Quietly building: more output than attention, for now.

+17 understory score output 88 · clout 71
Aug 2025 Jul 2026
  • output, commits & releases
  • clout, star velocity

output = commits & releases; clout = star velocity, both 0 to 100 monthly indices; the gap where output runs above clout is the understory. The understory →

covered in

  • PR#7 2025-04-30 on the radar

    A Comprehensive End-to-End Speech Recognition Toolkit

similar projects

compare these →
  • 🎙️ meetily

    Rust

    Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes

    29k ACTIVE
  • 🎙️ Real-Time-Voice-Cloning

    3× the stars

    Clone a voice in 5 seconds to generate arbitrary speech in real-time

    60k ACTIVE
  • 👁️ supervision

    2.5× the stars

    We write your reusable computer vision tools. 💜

    49k ACTIVE

comments

Sign in with GitHub to add your blip on FunASR.

loading comments…