Built for voice agents that can't wait 600ms

Google Gemini: Gemini 3.5 Transcribe

gemini/gemini-3-5-transcribe
Speech to textactive

Google's dedicated speech-to-text model (launched 2026-08-26): pre-recorded audio transcription via the Interactions API with 85+ auto-detected languages, speaker diarization (up to 8 speakers), word-level timestamps, custom vocabulary, and verbatim or smart (disfluency-cleaned, formatted) transcripts.

ProviderSlugPriceLatencyUptime