Audio-Video Understanding176 / 372 projects
Audio-Visual model leaderboard featuring the hottest audio and video understanding AI projects on GitHub for speech recognition and media analysis.
Cross-platform, customizable ML solutions for live and streaming media.
Deezer source separation library including pretrained models.
SoftVC VITS Singing Voice Conversion
DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Meta's generative audio AI. MusicGen for music, AudioGen for sound effects, and EnCodec for compression.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
kaldi-asr/kaldi is the official location of the Kaldi project.
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
A PyTorch-based Speech Toolkit
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
End-to-End Speech Processing Toolkit
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Speech recognition module for Python, supporting several engines and APIs, online and offline.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
Code for the paper "Jukebox: A Generative Model for Music"
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Facebook AI Research's Automatic Speech Recognition Toolkit
On-device Speech AI for Apple Silicon
Microsoft's unified speech-text model. TTS, ASR, voice conversion, and speech translation.
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Production First and Production Ready End-to-End Speech Recognition Toolkit
Muzic: Music Understanding and Generation with Artificial Intelligence
On-device wake word detection powered by deep learning
A small speech recognizer
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
Japanese text-to-speech engine with many character voices. Open-source, used for content creation.
Stable diffusion for real-time music generation
A simple, high-quality voice conversion tool focused on ease of use and performance.
Lingvo
Core Engine of Singing Voice Conversion & Singing Voice Clone
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.
:microphone: React Native Voice Recognition library for iOS and Android (Online and Offline Support)
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
An AI for Music Generation
中文语音识别; Mandarin Automatic Speech Recognition;
On-device, real-time multimodal AI with features similar to GPT-Live
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
Open-Source Large Vocabulary Continuous Speech Recognition Engine
:unlock: Lip Reading - Cross Audio-Visual Recognition using 3D Architectures
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.
DELTA is a deep learning based natural language and speech processing platform. LF AI & DATA Projects: https://lfaidata.foundation/projects/delta/
Toolkit for efficient experimentation with Speech Recognition, Text2Speech and NLP
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
SALMONN family: A suite of advanced multi-modal LLMs
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)
A fundamental toolkit designed for music, song, and audio generation
A fundamental toolkit designed for music, song, and audio generation
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
SincNet is a neural architecture for efficiently processing raw audio samples.
Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
Offline speech recognition for Android with Vosk library.
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
an open-source implementation of sequence-to-sequence based speech processing engine
A lightweight, simple-to-use, RNN wake word listener
Implementation of the Wave-U-Net for audio source separation
Whisper.net. Speech to text made simple using Whisper Models
Espresso: A Fast End-to-End Neural Speech Recognition Toolkit
Optimized Whisper models for streaming and on-device use
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
Voice activity detection (VAD) toolkit including DNN, bDNN, LSTM and ACAM based VAD. We also provide our directly recorded dataset.
Examples of how to use or integrate DeepSpeech
中文语音识别
GLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parameters
Connectionist Temporal Classification (CTC) decoding algorithms: best path, beam search, lexicon search, prefix search, and token passing. Implemented in Python.
The official repository of the Eesen project
基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。
A real-time silent speech recognition tool.
Allosaurus is a pretrained universal phone recognizer for more than 2000 languages
Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。
Private and on-device speech recognition keyboard and service for Android.
Foundational Model for Speech Recognition Tasks
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
On-device Speech-to-Intent engine powered by deep learning
👁 👂 Smart media tagging for Nextcloud: recognizes faces, objects, landscapes, music genres
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
:speech_balloon: An On-Premises, Streaming Speech Recognition System
On-device streaming speech-to-text engine powered by deep learning
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
Open-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.
End-to-end ASR/LM implementation with PyTorch
Connectionist Temporal Classification (CTC) decoder with dictionary and language model.
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
这是一个用C++实现ASR推理的项目,它依赖很少,安装也很简单,推理速度很快,在树莓派4B等ARM平台也可以流畅的运行。 支持的模型是由Google的Transformer模型中优化而来,数据集是开源wenetspeech(10000+小时)或阿里私有数据集(60000+小时), 所以识别效果也很好,可以媲美许多商用的ASR软件。
A speech recognition framework designed for SwiftUI.
Automatic Speech Recognition(ASR), Text-To-Speech(TTS) engine. 中英语音识别、多角色语音合成,支持多语言,准确率高
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
Phonetisaurus G2P
VOSK Speech Recognition Toolkit
Dataset and code of GTSinger(NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
libfaceid is a research framework for prototyping of face recognition solutions. It seamlessly integrates multiple detection, recognition and liveness models w/ speech synthesis and speech recognition.
UniSpeech - Large Scale Self-Supervised Learning for Speech
On-device speech-to-text engine powered by deep learning
Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
🇺🇦 Speech Recognition & Synthesis for Ukrainian
Arabic speech recognition, classification and text-to-speech.
This repository has implementation for "Neural Voice Cloning With Few Samples"
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
speech to text with self-supervised learning based on wav2vec 2.0 framework
A live speech recognition using Facebooks wav2vec 2.0 model.
Video to Text: Natural language description generator for some given video. [Video Captioning]
Kaldi-based Korean ASR (한국어 음성인식) open-source project
music generation with masked transformers!
Local windowed attention multi-instrumental music transformer for supervised music generation
Deep Learning based Automatic Speech Recognition with attention for the Nvidia Jetson.
The BEST music separation model with help of A.I. ... to my ears ! 👂👂
The official implementation of Theme Transformer. A Theme-based music generation. IEEE TMM
ZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.
Library of state-of-the-art models (PyTorch) for NLP tasks
Official implementation for Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
SummerAsr is a C++-based local Chinese speech recognizer that can be compiled independently with minimal dependencies.
Pytorch implementation of Noisy Student Training for Automatic Speech Recognition and Automatic Pronunciation Error Detection problem
Speech-to-text based on wav2letter built for transfer learning
Cross-lingual Voice Conversion
Implementation of the paper "Self-supervised Learning with Random-projection Quantizer for Speech Recognition" in Pytorch.
WavEncoder is a Python library for encoding audio signals, transforms for audio augmentation, and training audio classification models with PyTorch backend.
Traditional ASR (Signal & Cepstral Analysis, DTW, HMM) & DNNs (Custom Models + DeepSpeech) on Indian Accent Speech
Tensorflow implementation of "Listen, Attend and Spell" authored by William Chan. This project utilizes input pipeline and estimator API of Tensorflow, which makes the training and evaluation truly end-to-end.
Repository containing experimentation platform on how to train, infer on wav2vec2 models.
openai/whisper + extra features
Speech Recognition model based off of FAIR research paper built using Pytorch.
A multilingual ASR model that can recognize ten Turkic languages—Azerbaijani, Bashkir, Chuvash, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Uyghur, and Uzbek.
An automatic speech recognition API
Open source offline speech recognition for Android using Mozilla's DeepSpeech in Termux
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
A MXNet implementation of Baidu's DeepSpeech architecture
speech-to-text in pytorch
Emotions recognition from audio and text files (only russian language)
Transcribing Speech with Multinomial Diffusion, training code and models.
Fastest Open Source TTS Model
[ICASSP 2020] CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition (A PyTorch implementation of Continuous Integrate-and-Fire mechanism).
Converts spoken words into text form.
Drax: Speech Recognition with Discrete Flow Matching
Keras (tensorflow) implementation of SincNet (Mirco Ravanelli, Yoshua Bengio - https://github.com/mravanelli/SincNet)
Multi-speaker Speech Synthesis Using VITS(KO, JA, EN, ZH)
Multilingual Speech Recognition for Indonesian Languages
Conv-LSTM-CTC speech recognition network (end-to-end), written in TensorFlow.
mixlingual speech recognition system; hybrid (GMM+NNet) model; Kaldi + Keras
A Windows client SDK and demo software for the ASRT speech recognition system.
🗣️ Speech recognition on Unity and Android without the annoying google popup!
Proof of concept app that demonstrates use of KeenASR SDK in ObjC. WE ARE HIRING: https://keenresearch.com/careers.html
ASRDeepspeech x Sakura-ML (English/Japanese) with deepspeech2 model in pytorch with support from Zakuro AI.
Code related to the Dutch instance and user groups of the KALDI speech recognition toolkit
Brasil TTS é um conjunto de sintetizadores de voz, em português do Brasil, que lê telas para portadores de deficiência visual. Transforma texto em áudio, permitindo que pessoas cegas ou com baixa visão tenham acesso ao conteúdo exibido na tela. Embora o principal público-alvo de sistemas de conversão texto-fala – como o Brasil TTS – seja formado por pessoas com deficiência visual, esse tipo de programa pode ser usado por pessoas com dislexia e outras dificuldades de leitura, pessoas com deficiência severa de fala, bem como por crianças pré-alfabetizadas. Além de ser uma ferramenta de tecnologia assistiva, sintetizadores de voz podem ter ainda aplicações pedagógicas e de entretenimento.
Taiwan Tongues ASR CE 是一個開源語音辨識(Automatic Speech Recognition, ASR)模型專案,專為台灣多元語言環境設計。 本模型支援 國語、台語、客語與英語,提供本地多語混合語音辨識,讓開發者與資訊服務業者可運用此開源模型進行 ASR 模型訓練、微調與發展在地化應用,以低成本、高效率進行 ASR 語音應用落地與智慧服務創新。 本專案為數位發展部數位產業署「114年數位產業跨域軟體基盤系統建置案」之實證成果之一,旨在推動台灣語音技術開源生態,協助資訊服務業者強化智慧應用能量,落實在地 AI 技術自主發展,由台灣大哥大執行與維護。
Repository for the paper "Combining audio control and style transfer using latent diffusion", accepted at ISMIR 2024
Syn.Speech is a flexible speaker independent continuous speech recognition engine for Mono and .NET framework
Official code for Interspeech 2023 paper "Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering"
On-device VAD / streaming STT / TTS / diarization in C++17 (ONNX + LiteRT) with a voice-agent pipeline. Linux, Windows, Android.
speech-enhacement
Wav2vec 2.0 Self-Supervised Pretraining
Implementation of the paper "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations" in Pytorch.
Source code of the model used in Tensorflow Speech Recognition Challenge (https://www.kaggle.com/c/tensorflow-speech-recognition-challenge). The solution ranked in top 5% in private leaderboard.
Распознавание речи русского языка используя Tensorflow, обучаясь на базе Voxforge
Official repository for Big-Little Net
This repo is text to speech with learnable audio encoder without alignment with transcript reference
Text prompt steered synthetic audio generators
A mini, simple, and fast end-to-end automatic speech recognition toolkit.
Speech Recognition or Wake Word detection demo, developed using Maixduino framework and PlatfomIO, to run on K210 MCU on Sipeed's Maix dev board
[ISMIR 2022] Transfer Learning of wav2vec 2.0 for Automatic Lyric Transcription
Source: GitHub API · Curated · Realtime