Explore AI models on Synexa — 73 models, one API

Audio & Speech

Music, speech and sound effects — plus transcription back to text.

Sort
suno-latest

suno/suno-latest

Suno v5.5 generates full-length, high-fidelity AI songs with vocals and instrumentation from a single text prompt, producing 44.1 kHz stereo MP3 output.

Text to Audio$0.10/ request
ace-step

ace-step/ace-step

ACE-Step generates music with sung lyrics from a list of genre tags, extremely cheaply.

Text to Audio$0.0002/ second
kokoro

hexgrad/kokoro

Kokoro is a lightweight American-English text-to-speech model that is fast and very cheap to run.

Text to Audio$0.02/ 1k characters
lyria-2

google/lyria-2

Lyria 2 generates 30 seconds of 48kHz music from a descriptive text prompt.

Text to Audio$0.10/ request
mimo-tts

xiaomi/mimo-tts

Xiaomi MiMo V2.5 text-to-speech. Speaks text with one of 9 built-in voices, clones a voice from a reference clip, or invents a new voice from a written description.

Text to Audio$0.02/ request
music-3

minimax/music-3

MiniMax Music 3 writes and performs a complete song from lyrics and a style description.

Text to Audio$0.002/ second
music-generator

cassetteai/music-generator

A very fast instrumental music generator: a 30-second sample in under two seconds.

Text to Audio$0.02/ minute
seed-audio-1.0

bytedance/seed-audio-1.0

Seed Audio 1.0 generates natural speech, and can clone a voice from short reference clips.

Text to Audio$0.1875/ minute
sound-generation

elevenlabs/sound-generation

ElevenLabs sound effects. Turns a written description into a sound effect of up to 30 seconds — footsteps, weather, impacts, UI stings.

Text to Audio$0.05/ request
stable-audio-2.5

stability-ai/stable-audio-2.5

Stable Audio 2.5 generates music and sound effects from a text description.

Text to Audio$0.20/ request
tts

elevenlabs/tts

ElevenLabs text to speech. Reads text aloud in one of 21 built-in voices, across 29 languages on the default model and up to 74 on Eleven v3.

Text to Audio$0.05/ request
tts-v3

elevenlabs/tts-v3

Eleven v3 reads text aloud with expressive, emotionally aware delivery across dozens of languages.

Text to Audio$0.10/ 1k characters