AI Wisdom
๐ŸŽ™๏ธ

Speech & Audio

Speech recognition, text-to-speech, voice cloning, and AI music generation.

Production ยท 6Stable ยท 3Experimental ยท 211 total
โ† All categories
5/5

OpenAI's hosted speech-to-text model โ€” more accurate than Whisper via the API

OpenAI's own successor to Whisper for API use โ€” built on a GPT-4o-class audio model with materially lower word-error rates, especially on accented speech and noisy audio. Streaming supported. Whisper remains the right call when you need to self-host; this is the right call when you're already calling the API.

Proprietary
5/5

OpenAI's open-source speech recognition model โ€” still the gold standard for self-hosted ASR

Best open-source ASR model by far. 99 languages, timestamps, speaker diarization via community extensions. Self-host for free โ€” OpenAI's own GPT-4o Transcribe now leads on raw accuracy for API users, but nothing open-weight beats Whisper for offline or on-prem transcription.

Open Source

ElevenLabs

Production
5/5

Best-in-class text-to-speech with voice cloning and dubbing

Most natural-sounding TTS available. Voice cloning from 30 seconds of audio. Multilingual dubbing preserves emotion and cadence. API-first with generous free tier.

Proprietary

OpenAI TTS

Production
4/5

OpenAI's simple API for high-quality speech synthesis

Clean, natural-sounding voices with a simple API. The newer gpt-4o-mini-tts model adds instructable delivery โ€” you can prompt for tone and style ('speak like a calm customer support agent'), not just pick a voice. Best for developers wanting quick TTS integration without complexity.

Proprietary

Deepgram

Production
4/5

Real-time speech-to-text API with sub-300ms latency

Fastest production ASR API โ€” under 300ms latency for real-time use. Nova-3 model improved on Nova-2 with multilingual code-switching (handling a language switch mid-sentence). Best for live transcription, call centres, and real-time applications.

Proprietary

AssemblyAI

Production
4/5

Transcription API with built-in audio intelligence features

Excellent transcription with audio intelligence layered on top โ€” sentiment analysis, topic detection, PII redaction, summarization. Universal-2 model handles accents and noise well.

Proprietary

Suno v4.5

Stable
4/5

AI music generation from text prompts with full song structure

Generate full songs with vocals, instruments, and structure from text descriptions. The 4.5 update improved audio fidelity and genre blending noticeably over v4. Quality approaching professional production. Game-changer for content creators, ads, and game audio.

Proprietary

Udio

Stable
4/5

AI music creation with fine-grained style and genre control

Best genre fidelity in AI music โ€” excels at specific styles from jazz to electronic. Higher audio quality than competitors. Good for musicians wanting AI-assisted composition.

Proprietary

Bark

Stable
3/5

Open-source TTS model with emotions, music, and sound effects

Unique open-source TTS that generates speech with laughing, singing, and sound effects. Quality less consistent than ElevenLabs but fully self-hostable. Great for creative applications.

Open Source

Kokoro TTS

Experimental
3/5

Lightweight 82M-parameter open-source text-to-speech model

Impressively natural speech from a tiny 82M model โ€” runs on CPU. Good for edge deployment and resource-constrained environments. Quality approaching much larger models.

Open Source

Coqui TTS

Experimental
3/5

Open-source TTS toolkit for training and deploying custom voices

Most flexible open-source TTS toolkit. Train custom voices on your own data. XTTS model supports 16 languages with voice cloning. Best for teams needing fully custom voice solutions.

Open Source