Built for voice agents that can't wait 600ms

ElevenLabs: Eleven v4

elevenlabs/eleven-v4
active tts 2 providers
Playground

ElevenLabs' most expressive text-to-speech model, with audio-tag direction, Stability and Similarity voice settings, and 85 documented languages.

AuthorElevenLabs
Languages85 languages
VoicesLoading...
Response formatsSynchronous: MP3 · WAV · PCM · Opus · µ-law · A-law
Streaming: MP3 · PCM · Opus · µ-law · A-law
Statusactive

Providers

Different providers can host the same model. Choose one provider when you need a fixed backend, or let routing select among them.

ProviderSlugPer 1k charactersLatencyiFor TTS, latency is measured as time to first byte of audio. For STT, latency is measured as time to first token.ModesUptime
ElevenLabs 14 params
elevenlabs
$0.08
—
Synchronous
Fal 1 params
fal
$0.08
—
Synchronous

Supported languages

85 languages
Afrikaans af
Arabic ar
Armenian hy
Assamese as
Asturian ast
Azerbaijani az
Belarusian be
Bengali bn
Bosnian bs
Bulgarian bg
Burmese my
Cantonese yue
Catalan ca
Cebuano ceb
Chichewa ny
Croatian hr
Czech cs
Danish da
Dutch nl
English en
Estonian et
Filipino fil
Finnish fi
French fr
Galician gl
Georgian ka
German de
Greek el
Gujarati gu
Hausa ha
Hebrew he
Hindi hi
Hungarian hu
Icelandic is
Indonesian id
Irish ga
Italian it
Japanese ja
Javanese jv
Kannada kn
Kazakh kk
Kirghiz ky
Korean ko
Latvian lv
Lingala ln
Lithuanian lt
Luxembourgish lb
Macedonian mk
Malay ms
Malayalam ml
Maltese mt
Mandarin Chinese zh
Maori mi
Marathi mr
Mongolian mn
Nepali ne
Norwegian no
Occitan oc
Odia or
Pashto ps
Persian fa
Polish pl
Portuguese pt
Punjabi pa
Romanian ro
Russian ru
Serbian sr
Sindhi sd
Slovak sk
Slovenian sl
Somali so
Spanish es
Swahili sw
Swedish sv
Tajik tg
Tamil ta
Telugu te
Thai th
Turkish tr
Ukrainian uk
Urdu ur
Uzbek uz
Vietnamese vi
Welsh cy
Yoruba yo

Voices

Features

Text to speech

Generates spoken audio from text input.

text_to_speech
Streaming

Streams synthesized audio incrementally over HTTP; realtime WebSocket streaming uses the separate Text to Dialogue WebSocket.

streaming
Synchronous

Synthesizes speech from fully-submitted text in a single synchronous request.

synchronous
Multilingual

Supports 85 languages.

multilingual
Expressive speech

Designed for expressive, natural, or emotionally rich generated speech.

expressive_speech

Code

Client
Mode
Language

ElevenLabs SDK streaming TTS

Routes the ElevenLabs SDK through allmodels and streams the audio response.

import { ElevenLabsClient } from "elevenlabs";

const client = new ElevenLabsClient({
  apiKey: process.env.ALLMODELS_API_KEY,
  baseUrl: "https://api.allmodels.io/el"
});

const stream = await client.textToSpeech.stream("voice_id", {
  text: "I'm sorry, Dave. I'm afraid I can't do that.",
  modelId: "elevenlabs/eleven-v4",
  outputFormat: "mp3_44100_128"
});

const reader = stream.getReader();
for (;;) {
  const { done, value } = await reader.read();
  if (done) break;
  // Write value to your response, file, or audio playback pipeline.
}