Qwen Alibaba Cloud: Qwen3-ASR-1.7B
qwen/qwen3-asr-1-7b
Speech to textactive
Qwen's 1.7B open-weight multilingual speech recognition model. Transcribes pre-recorded audio with word and segment timestamps and optional speaker diarization, and also runs as a low-latency streaming transcriber over an OpenAI-Realtime-shaped WebSocket. The upstream checkpoint advertises 52 languages and dialects; a hosting provider may serve a narrower set, so the served list is recorded per route.
ProviderSlugPriceLatencyUptime
