TOPIC GUIDE
Voice AI: speech-to-text, text-to-speech and voice agents
Compare ElevenLabs, Deepgram, AssemblyAI, Cartesia, Vapi and Retell AI by what you are building: transcription, speech or phone agents.
- products
- 17
- things to compare
- 3
- updated
- 2026-10-02
Voice products split into three layers: turning speech into text, turning text into speech, and agents that hold a conversation over the phone or web. Pick tools by layer, then test them on your own audio, accents and latency budget.
What to compare
- Test transcription on your own recordings, including accents, noise and domain vocabulary.
- Measure end-to-end latency for live agents, not just model speed.
- Compare per-minute pricing at your expected monthly volume, including provider costs passed through by agent platforms.
This is a focused selection, not an exhaustive ranking. See each profile for evidence and review scope.
Projects to explore
ElevenLabs
An AI voice generator and voice agents platform with thousands of voices in over 70 languages, plus APIs and SDKs.
Read the project profileDeepgram
Voice AI APIs for real-time and batch speech-to-text, text-to-speech and voice agents, built for enterprise scale.
Read the project profileAssemblyAI
Speech AI models and APIs to transcribe audio and extract insights, plus a voice agent API.
Read the project profileCartesia
A real-time text-to-speech API with expressive voices in 44 languages, built for AI agents and interactive apps.
Read the project profileVapi
A developer platform to build, test and deploy voice AI agents that handle phone and web calls.
Read the project profileRetell AI
A platform for deploying AI voice agents that answer and make phone calls, with pay-as-you-go pricing per minute.
Read the project profileMurf AI
An AI voice platform for studio voiceovers, dubbing, a text-to-speech API and conversational voice agents.
Read the project profileSpeechify
Text to speech that reads documents, PDFs, web pages and books aloud in natural voices, with voice typing and an AI assistant.
Read the project profileWispr Flow
Voice dictation that turns natural speech into polished text in any app, in 100+ languages, plus an AI meeting notetaker.
Read the project profileTurboScribe
Transcribes audio and video files to text in 98+ languages using Whisper, and translates transcripts or subtitles into 134+ languages.
Read the project profileSuperwhisper
AI voice-to-text that works in any app on macOS, Windows, iOS and Android, with offline transcription and custom modes.
Read the project profileBland
A platform for building, deploying and monitoring AI phone agents, aimed at regulated industries.
Read the project profilePipecat
An open-source Python framework for building voice and multimodal conversational agents, maintained by Daily.
Read the project profileLiveKit Agents
LiveKit's open-source framework for building real-time voice, video and physical AI agents in Python or Node.js.
Read the project profileMacWhisper
A Mac app that transcribes audio, video, meetings and system audio locally with Whisper, Parakeet and other models.
Read the project profilewhisper.cpp
A dependency-free C/C++ port of OpenAI's Whisper speech recognition that runs on Mac, Windows, Linux, iOS, Android, Raspberry Pi and the web.
Read the project profileChatterbox
Resemble AI's open-source text-to-speech models with voice cloning across many languages.
Read the project profile