Built for voice agents that can't wait 600ms

Inworld AI: Realtime TTS-2

inworld/inworld-tts-2
Text to speechactive

Inworld's flagship text-to-speech model: 200+ languages and locales, natural-language steering with persistent bracketed instruction tags, word/character timestamps with phoneme and viseme detail, instant and professional voice cloning, and a reported 100 ms P90 server-side time to first audio byte.

ProviderSlugPriceLatencyUptime