语音与数字人
语音识别、语音合成、数字人形象与口播视频
数据截至 8/1 18:48(构建快照,正在获取最新)
按仍在维护排序:先筛掉近 30 天没有提交的项目,再按热度排。 高 star 但早已停更的项目不会出现在这里。
- 1
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
- 2
Port of OpenAI's Whisper model in C/C++
- 3
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
- 4
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
- 5
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
- 6
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
- 7
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
- 8
Translate the video from one language to another and embed dubbing & subtitles.
- 9
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
- 10
🧠 Leon is your open-source personal assistant.
- 11
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
- 12
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
- 13
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
- 14
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
- 15
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
- 16
End-to-End Speech Processing Toolkit
- 17
Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.
- 18
Speech recognition module for Python, supporting several engines and APIs, online and offline.
- 19
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
- 20
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
- 21
💬 Speech recognition for your site
- 22
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
- 23
Facebook AI Research's Automatic Speech Recognition Toolkit
- 24
On-device Speech AI for Apple Silicon
- 25
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
- 26
Silero Models: pre-trained text-to-speech models made embarrassingly simple
- 27
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
- 28
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching
- 29
开源 AI 视频本地化工具:自动完成 YouTube/Bilibili 视频下载、字幕识别与翻译、语音克隆配音、音轨混合和字幕压制。
- 30
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
另有 43 个项目因近期无提交或缺少数据未列入。
语音与数字人要落到业务里,还差什么?
开源项目给的是能力,不是方案。数据怎么接、权限怎么管、上线后谁维护,这些才是落地的真正成本。