Submit projectSubmit

TOPIC GUIDE

Voice AI: speech-to-text, text-to-speech and voice agents

Compare ElevenLabs, Deepgram, AssemblyAI, Cartesia, Vapi and Retell AI by what you are building: transcription, speech or phone agents.

products
17
things to compare
3
updated
2026-10-02

Voice products split into three layers: turning speech into text, turning text into speech, and agents that hold a conversation over the phone or web. Pick tools by layer, then test them on your own audio, accents and latency budget.

What to compare

  1. Test transcription on your own recordings, including accents, noise and domain vocabulary.
  2. Measure end-to-end latency for live agents, not just model speed.
  3. Compare per-minute pricing at your expected monthly volume, including provider costs passed through by agent platforms.

This is a focused selection, not an exhaustive ranking. See each profile for evidence and review scope.

Projects to explore

ElevenLabs

An AI voice generator and voice agents platform with thousands of voices in over 70 languages, plus APIs and SDKs.

Read the project profile

Deepgram

Voice AI APIs for real-time and batch speech-to-text, text-to-speech and voice agents, built for enterprise scale.

Read the project profile

AssemblyAI

Speech AI models and APIs to transcribe audio and extract insights, plus a voice agent API.

Read the project profile

Cartesia

A real-time text-to-speech API with expressive voices in 44 languages, built for AI agents and interactive apps.

Read the project profile

Vapi

A developer platform to build, test and deploy voice AI agents that handle phone and web calls.

Read the project profile

Retell AI

A platform for deploying AI voice agents that answer and make phone calls, with pay-as-you-go pricing per minute.

Read the project profile

Murf AI

An AI voice platform for studio voiceovers, dubbing, a text-to-speech API and conversational voice agents.

Read the project profile

Speechify

Text to speech that reads documents, PDFs, web pages and books aloud in natural voices, with voice typing and an AI assistant.

Read the project profile

Wispr Flow

Voice dictation that turns natural speech into polished text in any app, in 100+ languages, plus an AI meeting notetaker.

Read the project profile

TurboScribe

Transcribes audio and video files to text in 98+ languages using Whisper, and translates transcripts or subtitles into 134+ languages.

Read the project profile

Superwhisper

AI voice-to-text that works in any app on macOS, Windows, iOS and Android, with offline transcription and custom modes.

Read the project profile

Bland

A platform for building, deploying and monitoring AI phone agents, aimed at regulated industries.

Read the project profile

Pipecat

An open-source Python framework for building voice and multimodal conversational agents, maintained by Daily.

Read the project profile

LiveKit Agents

LiveKit's open-source framework for building real-time voice, video and physical AI agents in Python or Node.js.

Read the project profile

MacWhisper

A Mac app that transcribes audio, video, meetings and system audio locally with Whisper, Parakeet and other models.

Read the project profile

whisper.cpp

A dependency-free C/C++ port of OpenAI's Whisper speech recognition that runs on Mac, Windows, Linux, iOS, Android, Raspberry Pi and the web.

Read the project profile

Chatterbox

Resemble AI's open-source text-to-speech models with voice cloning across many languages.

Read the project profile

Go a little deeper