Qwen3.8-LiveTranslate processes speech with an average latency of 2.3 seconds

30 September 202610 views

The Interleave architecture model recognizes 60 languages, speaks 29, and supports translation during conversations. The article covers reducing average latency, separating utterances by speaker, bilingual output, and access via the Alibaba Cloud and QwenCloud APIs.

Qwen3.8-LiveTranslate processes speech with an average latency of 2.3 seconds

Streaming Translation Latency

Qwen3.8-LiveTranslate listens to a speaker and provides a text and spoken translation while they are talking. Average delay is measured using the LAAL metric, which shows how far the translation lags behind the original speech.

LAAL metricValue
Before2.8 seconds
Now2.3 seconds
Changeapproximately 18% less

According to Qwen, the new Interleave architecture also improves accuracy in conveying meaning, fluency, and translation brevity. A practical criterion for choosing it is whether you need the translation before the speaker finishes talking.

Input and output

Qwen3.8-LiveTranslate accepts audio and optional images. Video frames can be sent along with speech.

The model outputs text and audio. This makes its format suitable for scenarios that require both a written translation and spoken output.

Languages and conversation features

Qwen3.8-LiveTranslate understands 60 languages and can provide spoken output in 29 of them. The remaining 31 are available for text translation only.

The model recognizes different speakers in multi-party speech and preserves their voices. It also displays the original text and translation in sync, and uses conversation history to translate names and terms consistently.

A criterion for choosing it is whether the required language pair supports spoken output and whether you need multi-speaker support.

Integration

Qwen3.8-LiveTranslate is available through the WebSocket Realtime API on Alibaba Cloud Model Studio and QwenCloud. The model ID is qwen3.8-livetranslate-flash-realtime.

Consider this model if your integration needs to use the WebSocket Realtime API and one of these services. The sources do not describe integration options beyond these platforms.

Frequently asked questions

Qwen3.8-LiveTranslate processes speech with an average latency of 2.3 seconds