Awesome AI Agent Stack
Voice, Speech & Audio
Speech-to-text, text-to-speech and voice agents.
Speech-to-text
- openai/whisper - Robust Speech Recognition via Large-Scale Weak Supervision.
- ggml-org/whisper.cpp - Port of OpenAI's Whisper model in C/C++.
- SYSTRAN/faster-whisper - Faster Whisper transcription with CTranslate2.
- m-bain/whisperX - WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization).
- NVIDIA-NeMo/Speech - NVIDIA's toolkit for building production ASR, TTS, and conversational AI.
- modelscope/FunASR - Alibaba's industrial-grade speech recognition toolkit for production ASR.
- facebookresearch/seamless_communication - Foundational models for state-of-the-art speech and text translation.
- QwenAudio/SenseVoice - Multilingual ASR with language ID, emotion and audio-event detection.
- k2-fsa/sherpa-onnx - Offline ASR, TTS, diarization, VAD and enhancement via ONNX Runtime.
- k2-fsa/sherpa - Speech-to-text server framework built on next-gen Kaldi.
- speechbrain/speechbrain - PyTorch toolkit for ASR, diarization and speech processing.
- alphacep/vosk-api - Offline speech recognition for mobile, Raspberry Pi and servers.
- espnet/espnet - End-to-end speech processing toolkit for ASR, TTS and translation.
- PaddlePaddle/PaddleSpeech - Speech toolkit: streaming ASR/TTS, punctuation, speaker verification.
- argmaxinc/argmax-oss-swift - On-device speech AI for Apple Silicon (WhisperKit).
- ufal/whisper_streaming - Real-time streaming Whisper transcription and translation.
- collabora/WhisperLive - A nearly-live implementation of OpenAI Whisper.
- speaches-ai/speaches - OpenAI-compatible speech-to-text and text-to-speech inference server.
- ahmetoner/whisper-asr-webservice - Whisper ASR exposed as a webservice API.
- KoljaB/RealtimeSTT - Low-latency speech-to-text with VAD and wake-word activation.
- Vaibhavs10/insanely-fast-whisper - Whisper transcription accelerated with transformers and optimum.
- linto-ai/whisper-timestamped - Multilingual ASR with word-level timestamps and confidence scores.
- HaujetZhao/CapsWriter-Offline - Offline hotkey-driven dictation app with high accuracy and low latency.
- moonshine-ai/moonshine - Very low-latency speech-to-text for voice agents and interfaces.
- Picovoice/cheetah - On-device streaming speech-to-text engine.
- Picovoice/leopard - On-device speech-to-text engine for private local transcription.
- kyutai-labs/hibiki - Simultaneous speech translation that adapts its flow as you speak.
- wenet-e2e/wenet - Production-first, production-ready end-to-end speech recognition toolkit.
- k2-fsa/icefall - Recipes and training pipelines for next-gen Kaldi speech models.
- QuentinFuxa/WhisperLiveKit - Real-time local speech-to-text with streaming ASR and speaker diarization.
- Blaizzy/mlx-audio - Text-to-speech, speech-to-text and speech-to-speech library built on Apple MLX.
- thewh1teagle/vibe - Offline desktop app for transcribing audio and video with Whisper.
- MahmoudAshraf97/whisper-diarization - Whisper-based speech recognition with speaker diarization.
- FluidInference/FluidAudio - CoreML speech-to-text, text-to-speech and diarization models for Apple apps.
- jhj0517/Whisper-WebUI - Web UI for generating and translating subtitles with Whisper models.
- cjpais/Handy - Offline, extensible speech-to-text desktop app.
- zachlatta/freeflow - Free, fast dictation app; an open alternative to Wispr Flow.
Text-to-speech & voice cloning
Voice agents
Audio processing & analysis
Telephony & SIP