OpenAI text to speech
Generate speech with the official OpenAI JavaScript or Python SDK.
Set the OpenAI client's base URL to https://api.allmodels.io/oai and use your AllModels API key.
Generate an audio file
https://api.allmodels.io/oai/audio/speechOpenAIimport { writeFile } from "node:fs/promises";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.ALLMODELS_API_KEY,
baseURL: "https://api.allmodels.io/oai"
});
const audio = await client.audio.speech.create({
model: "elevenlabs/eleven-turbo-v2-5",
voice: "pNInz6obpgDQGcFmaJgB",
input: "Hello from AllModels",
response_format: "mp3"
});
await writeFile("speech.mp3", Buffer.from(await audio.arrayBuffer()));Streaming
Start playing or processing audio before the full response is ready.
import { createWriteStream } from "node:fs";
import { Readable } from "node:stream";
const response = await client.audio.speech.create({
model: "grok/grok-tts",
voice: "eve",
input: "This response starts playing before generation is complete.",
response_format: "pcm",
stream_format: "audio"
});
Readable.fromWeb(response.body).pipe(createWriteStream("speech.pcm"));Use stream_format: "audio" for audio bytes or stream_format: "sse" for OpenAI speech events. Support varies by model and provider.
Formats and controls
| Field | Description |
|---|---|
model | Model name or accepted alias. |
voice | Provider voice ID string or { "id": "..." } object. |
input | Text to synthesize. See the API reference for the maximum length. |
response_format | mp3, opus, aac, flac, wav, or pcm. |
speed | Speech rate where supported. |
instructions | Free-form delivery direction, up to 4096 characters. Honored on the bindings listed below; refused elsewhere. |
stream_format | Raw audio bytes or sse events. |
provider | AllModels provider preferences. |
provider_options | Provider-specific controls. |
AllModels does not convert audio between formats. If the selected model cannot produce the requested format, the API returns invalid_response_format.
Delivery instructions
instructions carries free-form direction — tone, pacing, character — of at most 4096 characters. Few models take one natively, so AllModels renders it into whatever mechanism the selected binding documents, and only where a live probe showed the direction actually reaching the audio:
| Binding | How the direction is rendered |
|---|---|
openai/gpt-4o-mini-tts | Forwarded unchanged in OpenAI's own instruction field. |
gemini/gemini-2.5-flash-preview-tts, gemini/gemini-2.5-pro-preview-tts, gemini/gemini-3.1-flash-tts-preview | Carried as its own section of the model's text prompt. |
gemini/gemini-3.8-flash-tts, gemini/gemini-3.8-flash-lite-tts | Sent in Google's dedicated speech_metadata.style field; input is spoken verbatim, so do not write directions into it. |
groq/canopylabs/orpheus-v1-english | Sanitized and prepended as one leading bracketed direction. |
fal/fal-ai/elevenlabs/tts/eleven-v3 | The same; this endpoint serves synchronous requests only. |
grok/grok-tts | Mapped onto Grok's own inline and wrapping tags. Synchronous requests only — this route honors it, the streaming TTS WebSocket refuses it. |
minimax/speech-2.8-hd, minimax/speech-2.8-turbo | Mapped onto MiniMax interjection markers leading the utterance. |
fal/fal-ai/minimax/speech-2.8-hd, fal/fal-ai/minimax/speech-2.8-turbo | The same markers; these endpoints serve synchronous requests only. |
The last three rows are fixed-tag bindings. They render a small verified vocabulary rather than arbitrary prose, so whether a direction is honored depends on its wording: "laugh, then whisper" is honored on grok/grok-tts and refused on MiniMax, which has no whisper tag. A direction the binding's vocabulary does not cover completely is refused rather than rendered in part.
Check the instructions capability on GET /v1/models for the current answer per binding — see Reading the instructions capability below.
Requests carrying an instructions value are routed only to bindings that honor it. When none can — including a pinned provider that cannot, and a fixed-tag binding whose vocabulary does not cover the wording — the request fails with 422 unsupported_parameter on instructions rather than returning speech that ignores the direction. Inline cues written directly into input are unaffected: they pass through to the provider verbatim, on any model.
One caution when writing inline cues yourself: not every model validates its markers, and we probed each binding's behavior. Some fail safe by stripping a marker they do not recognize (Grok, Fish, ElevenLabs v3, Cartesia, OpenAI tts-1, Soniox), and the Gemini 2.5 and 3.1 models interpret bracket tags as delivery direction (the 3.8 models use <angle> tags instead). But several speak unrecognized markers aloud — a typo'd or unsupported cue like (clears throat) or [weeping] comes back as audible words in the speech: both MiniMax families, Together's Orpheus and Kokoro, Deepgram Aura-2, the ElevenLabs v2 family, and — intermittently — OpenAI's gpt-4o-mini-tts itself. On those models, prefer the instructions parameter where it is honored: AllModels only ever emits tags it has verified on that binding.
Reading the instructions capability
Every TTS binding on GET /v1/models carries an instructions object. Coding agents should read it before deciding how to express delivery:
"instructions": {
"mode": "inline-freeform",
"honored": false,
"docs": "https://docs.fish.audio/developer-guide/core-features/emotions",
"skill": "https://api.allmodels.io/skills/fish.md"
}mode— how this model accepts delivery direction (table below).honored—truemeans AllModels translates theinstructionsparameter for this binding.skill— a markdown guide we host for this binding's provider on how to steer delivery withinstructionsthrough AllModels: which models honor it, the tags each accepts, and example requests. It is served from the same host you called, and its example requests target that host. Always present.docs— the provider's own guide to emotions, audio tags and pacing. Absent when the provider publishes none.
Coding agents should fetch skill first, since it describes how to steer delivery through AllModels. Then read docs for depth on the provider's own vocabulary.
Each skill's frontmatter carries an updated date (when it was last checked against the provider's docs, also sent as the Last-Modified header) and stale_after_days. Once a skill is older than that, take the provider's current pass-through tag list from docs instead. The vocabularies AllModels verifies for honored inline-fixed bindings stay authoritative either way, because AllModels enforces them.
mode | What the model accepts | What to do |
|---|---|---|
native | A dedicated instruction field | Send instructions, e.g. "speak warmly, slowly". |
prompt | Direction inside the model's text prompt | Send instructions, or follow docs to write the direction yourself. |
inline-freeform | Open-ended cues in brackets inside the text | Put cues in input: "[whispers softly] The numbers look great." |
inline-fixed | Only a fixed set of documented tags | Use only tags listed in docs: "(laughs) The numbers look great." |
structured | Typed controls such as speed or emotion, no prose | Set the provider's options; do not put cues in the text. |
unsupported | No delivery control | Send plain text. |
When honored is true, prefer the instructions parameter: AllModels only emits tags it has verified on that binding. When writing cues into input yourself, stick to what docs lists — as the caution above explains, some models speak unrecognized markers aloud.
Select a provider
Use provider preferences to choose or prioritize providers:
await client.audio.speech.create({
model: "elevenlabs/eleven-turbo-v2-5",
voice: "pNInz6obpgDQGcFmaJgB",
input: "Route this through Fal.",
response_format: "mp3",
provider: { only: ["fal"], allow_fallbacks: false },
provider_options: { speed: 1.1 }
});See Providers for all provider options.
Handle errors
The OpenAI SDK raises an API status error when a request fails. Check the error code and status before retrying:
import OpenAI from "openai";
try {
await client.audio.speech.create({
model: "elevenlabs/eleven-turbo-v2-5",
voice: "pNInz6obpgDQGcFmaJgB",
input: "Hello from AllModels",
response_format: "mp3"
});
} catch (error) {
if (error instanceof OpenAI.APIError) {
console.error(error.status, error.error);
if (error.status === 429 || error.status >= 500) {
console.error("Retry with bounded exponential backoff");
}
} else {
throw error;
}
}