语音与数字人

语音识别、语音合成、数字人形象与口播视频

数据截至 8/1 18:48(构建快照,正在获取最新)

仍在维护排序:先筛掉近 30 天没有提交的项目,再按热度排。 高 star 但早已停更的项目不会出现在这里。

  1. 1
    GPT-SoVITSGitHub语音与数字人MIT

    1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

    60.3kstar10 天前提交Python
  2. 2
    whisper.cppGitHub语音与数字人MIT

    Port of OpenAI's Whisper model in C/C++

    52.5kstar1 天前提交C++
  3. 3
    VoxCPMGitHub语音与数字人Apache-2.0

    VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

    34.7kstar24 天前提交Python
  4. 4
    whisperXGitHub语音与数字人BSD-2-Clause

    WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

    23.4kstar19 天前提交Python
  5. 5
    index-ttsGitHub语音与数字人

    An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

    22.3kstar17 天前提交Python
  6. 6
    FunASRGitHub语音与数字人MIT

    Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

    19.6kstar1 天前提交Python
  7. 7
    pot-desktopGitHub语音与数字人GPL-3.0

    🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.

    19.1kstar27 天前提交JavaScript
  8. 8
    pyvideotransGitHub语音与数字人GPL-3.0

    Translate the video from one language to another and embed dubbing & subtitles.

    18.5kstar1 天前提交Python
  9. 9
    SpeechGitHub语音与数字人Apache-2.0

    A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

    17.8kstar今天有提交Python
  10. 10
    leonGitHub语音与数字人MIT

    🧠 Leon is your open-source personal assistant.

    17.4kstar今天有提交TypeScript
  11. 11
    vosk-apiGitHub语音与数字人Apache-2.0

    Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

    15.0kstar29 天前提交Jupyter Notebook
  12. 12
    sherpa-onnxGitHub语音与数字人Apache-2.0

    Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

    13.9kstar1 天前提交C++
  13. 13
    supertonicGitHub语音与数字人MIT

    Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

    13.6kstar8 天前提交Swift
  14. 14
    PaddleSpeechGitHub语音与数字人Apache-2.0

    Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

    12.7kstar11 天前提交Python
  15. 15
    voice-proGitHub语音与数字人GPL-3.0

    Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

    11.5kstar19 天前提交Python
  16. 16
    espnetGitHub语音与数字人Apache-2.0

    End-to-End Speech Processing Toolkit

    9.9kstar3 天前提交Python
  17. 17
    OmniVoice-StudioGitHub语音与数字人AGPL-3.0

    Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.

    9.4kstar1 天前提交Python
  18. 18
    speech_recognitionGitHub语音与数字人BSD-3-Clause

    Speech recognition module for Python, supporting several engines and APIs, online and offline.

    9.0kstar今天有提交Python
  19. 19
    SenseVoiceGitHub语音与数字人MIT

    Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

    9.0kstar4 天前提交C
  20. 20
    mlx-audioGitHub语音与数字人MIT

    A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

    7.7kstar今天有提交Python
  21. 21
    annyangGitHub语音与数字人MIT

    💬 Speech recognition for your site

    6.8kstar17 天前提交TypeScript
  22. 22
    espeak-ngGitHub语音与数字人GPL-3.0

    eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.

    6.7kstar今天有提交C
  23. 23
    wav2letterGitHub语音与数字人

    Facebook AI Research's Automatic Speech Recognition Toolkit

    6.4kstar18 天前提交C++
  24. 24
    argmax-oss-swiftGitHub语音与数字人MIT

    On-device Speech AI for Apple Silicon

    6.3kstar今天有提交Swift
  25. 25
    FunClipGitHub语音与数字人MIT

    FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.

    6.1kstar2 天前提交Python
  26. 26
    silero-modelsGitHub语音与数字人

    Silero Models: pre-trained text-to-speech models made embarrassingly simple

    6.0kstar今天有提交Jupyter Notebook
  27. 27
    abogenGitHub语音与数字人MIT

    Generate audiobooks from EPUBs, PDFs and text with synchronized captions.

    5.5kstar1 天前提交Python
  28. 28
    Kokoro-FastAPIGitHub语音与数字人Apache-2.0

    Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model w/multiplatform CPU, AMD, NVIDIA GPU PyTorch support, handling, and auto-stitching

    5.3kstar今天有提交Python
  29. 29
    YouDub-webuiGitHub语音与数字人Apache-2.0

    开源 AI 视频本地化工具:自动完成 YouTube/Bilibili 视频下载、字幕识别与翻译、语音克隆配音、音轨混合和字幕压制。

    5.2kstar17 天前提交Python
  30. 30
    dograhGitHub语音与数字人BSD-2-Clause

    Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.

    5.1kstar今天有提交Python

另有 43 个项目因近期无提交或缺少数据未列入。

语音与数字人要落到业务里,还差什么?

开源项目给的是能力,不是方案。数据怎么接、权限怎么管、上线后谁维护,这些才是落地的真正成本。