immich-app/immich
High performance self-hosted photo and video management solution.
// category
Speech, audio, image, and video: text-to-speech, transcription, OCR, vision, diffusion, and multimodal models.
These are our picks, not a scrape of GitHub trending: every repo here is one we covered in Repository Radar because we thought it worth tracking. Ranked by stars, with live metrics and a get-it block on each.
21 repos in this category. Browse or filter the full archive → · Compare all categories →
21 repos, our 5th-largest category (Agents & Orchestration leads with 92). Combined they're up 90% since we covered them, 5th-fastest-growing of our 9 categories. They hold 683k stars, 2k contributors and 1M weekly downloads between them. Strength runs broad, not top-heavy: the median repo (20k stars) sits close to the 33k average. 9 are breaking out and 20 still actively shipping, and Python leads on language.
total
per repo
immich-app/immich
High performance self-hosted photo and video management solution.
ruvnet/RuView
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection, all without a single pixel of video.
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
CorentinJ/Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
microsoft/VibeVoice
Open-Source Frontier Voice AI
roboflow/supervision
We write your reusable computer vision tools. 💜
jamiepine/voicebox
The open-source AI voice studio. Clone, dictate, create.
calesthio/OpenMontage
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
Zackriya-Solutions/meetily
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes
resemble-ai/chatterbox
SoTA open-source TTS
DrewThomasson/ebook2audiobook
Generate audiobooks from e-books, voice cloning & 1158+ languages!
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
browser-use/video-use
Edit videos with coding agents
KittenML/KittenTTS
State-of-the-art TTS model under 25MB 😻
KoljaB/RealtimeSTT
A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.
dograh-hq/dograh
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
VectorSpaceLab/OmniGen2
OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
Saiyan-World/goku
[CVPR2025 Highlight] Video Generation Foundation Models: https://saiyan-world.github.io/goku/
OmniSVG/OmniSVG
[NeurIPS 2025] OmniSVG is the first family of end-to-end multimodal SVG generators that leverage pre-trained Vision-Language Models (VLMs), capable of generating complex and detailed SVGs, from simple icons to intricate anime characters.
PrunaAI/pruna
Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.
magnitudedev/magnitude
Magnitude coding agent