Stormin' The Castle

Qwen3-TTSPerformanceInference

Can an RTX 5090 actually hit Qwen3-TTS's 97 ms latency?

by John Robinson @johnrobinsn

Qwen TTS Latency

TL;DR — Yes. On a stock RTX 5090 I measured median TTFA 47-51 ms (first-chunk delivery time) for Qwen3-TTS with the official vLLM-Omni streaming pipeline — about half of Qwen's "as low as 97 ms" headline. A sibling benchmark using transformers-native loading saw ~2.3 s TTFA on the same hardware — a ~50× gap, entirely attributable to the inference architecture. I also wired up precomputed-x-vector voice cloning (Base model) that pays the same TTFA as the predefined-speaker path (~59-62 ms across five voices). Repo + runnable samples: https://github.com/johnrobinsn/qwen3-tts-stream-demo

This article is paired with a public code repo at https://github.com/johnrobinsn/qwen3-tts-stream-demo — together they give you a step-by-step recipe to reproduce these numbers on an RTX 5090 (or any Blackwell-class GPU), with every environment pin, deploy-config override, and non-obvious gotcha already debugged. The repo also ships an interactive CLI tool — ./repl.sh for the predefined speakers and ./clone_repl.sh for the voice-cloning path — that lets you type text and hear it streamed through your speakers in real time, not just as a batch benchmark. And if you'd rather listen before installing anything, the showcase/ directory in the repo has ready-to-play 24 kHz WAVs covering every voice, language, and clone I discuss below.

Worth saying up front: the Qwen team (Alibaba) published the 97 ms number but not the inference code to reproduce it — the model card and technical report describe the Dual-Track streaming architecture at a high level, but the actual pipeline wiring, deploy profile, and per-request flags that make the number land are left to the reader. The open-source vllm-omni project is where that pipeline actually lives, and the Qwen3-TTS streaming optimizations have been actively landing over the last few weeks — the single-stage fused profile targeted specifically at first-audio latency went into main on 2026-10-02 (commit e84014df), the same day the first draft of this post was being written. That's what made this reproduction possible.

Quick backgrounder on Qwen3-TTS

For readers coming in cold: Qwen3-TTS is a family of open-weights neural TTS models from Alibaba's Qwen team — the same group behind the Qwen large-language-model line. All weights ship on Hugging Face under Qwen/Qwen3-TTS-12Hz-*. Highlights:

The number that didn't match

Qwen3-TTS launched with a claim worth paying attention to: "end-to-end synthesis latency as low as 97 ms", enabled by what they call a Dual-Track hybrid streaming generation architecture [1]. Everything about that description — streaming, dual-track, 97 ms — points at pipelined inference where multiple stages of the model work on different parts of the input concurrently so audio starts flowing before the whole utterance is "done."

A sibling benchmark on this same host measured something very different. Loading Qwen3-TTS with transformers.AutoModel.from_pretrained(..., attn_implementation="flash_attention_2") and running model.generate() gave an RTF of 1.85 (slower than real-time) and a first-chunk delay of 2.3–2.5 seconds on an RTX 5090. 50× off.

Numbers like that are almost never about the model. They're about the inference path.

The inference path matters more than the model

Qwen3-TTS factors into three stages:

  1. Text tokenizer — turns input text into token IDs.
  2. Talker — an autoregressive LLM that emits codec codes (not audio directly).
  3. Code2Wav — a lightweight codec decoder that converts those codes into PCM.

The transformers-native path runs these in sequence: tokenize → generate all codec codes → decode once → return the full WAV. First audio byte lands only after the last codec code is produced. For a 5-second utterance at RTF 1.85 that's about 2 seconds of staring at a progress bar.

The vLLM-Omni path runs the Talker and Code2Wav as separate processes connected by a shared-memory ring buffer. The Talker streams codec codes into the buffer as they're generated; Code2Wav consumes them in chunks and ships PCM back to the caller as fast as it can decode. By the time the Talker finishes the full sequence, most of the audio has already been played. The first chunk — specifically initial_codec_chunk_frames: 1 under the mrv2 deploy profile — gets priority routing on its own queue (talker_first_audio: true), so Code2Wav fires as soon as the Talker emits a single codec frame, instead of waiting for a steady-state 25-frame window to fill.

The piece that makes this surprising is that the Talker runs faster than real-time: emitting one codec code is a GPU inference step that takes tens of ms of wall-clock, not the 83 ms of audio content that one frame represents. That's why TTFA (~47 ms) can come in well under the audio duration of the first chunk (~83 ms). It's not that the model is faster than its own codec rate — it's that "first PCM delivered" and "first PCM finishes playing" are different clocks, and the vLLM-Omni pipeline optimises the former.

Hardware & software

GPU RTX 5090 (Blackwell sm_120, 32 GB)
Driver NVIDIA 595.84 (advertises CUDA 13.2)
CUDA toolkit 13.2 at /usr/local/cuda-13.2
Python 3.12.11 (uv-managed)
torch 2.13.0 + cu130
vLLM 0.30.0
vLLM-Omni 0.30.0
Flash-attn not installed — vLLM ships flashinfer + FA4-via-cutlass; vllm-omni uses fa3-fwd

Three constraints out of this list are worth calling out. First, vLLM 0.30 pins torch 2.13 + cu130 — you can't drop vllm-omni into an existing torch 2.8 / cu128 environment. Second, Blackwell sm_120 requires CUDA ≥ 12.9 to compile kernels, so the system's default nvcc (anaconda-shipped 12.8 on this host) will refuse to emit FlashInfer kernels until you prepend /usr/local/cuda-13.2/bin to PATH. Third, Debian/Ubuntu's system Python 3.12 carries a multi-arch pyconfig.h that breaks Triton's JIT — you have to use uv's prebuilt CPython. Each of these cost me an hour to debug; all three are documented in the repo README's troubleshooting table so you don't have to.

Methodology

For every measurement run:

  1. Warmup — one inference that doesn't count. The first pass through the Talker + Code2Wav triggers CUDA graph specialization and FlashInfer autotune (both one-time costs). The warmup prompt is deliberately long enough to exercise the full-size Code2Wav graph, not just the tiny first-chunk graph a one-word prompt would hit.
  2. Measurement runs — N inferences timed from AsyncOmni.generate submission to the first audio chunk being yielded to the Python iterator. This is TTFA. Also recorded: wall-clock to the final chunk, inter-chunk latency distribution, synthesized duration.
  3. Report — min / median / max TTFA across runs.

TTFA deliberately excludes OS-level audio stack latency (DAC startup, buffer draining). It's the model's clock, not the speaker's.

Results

CustomVoice (predefined speaker)

Model: Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice, speaker Ryan, English prompt.

Pass TTFA (ms) Chunks Total (ms) Inter-chunk p50 (ms)
warmup 781 5 1458 149
run 1 47 5 438 147
run 2 51 3 458 147
run 3 64 3 807 155

The first chunk is 1920 samples at 24 kHz = 80 ms of audio delivered in ~50 ms wall-clock. Steady-state chunks are 48 000 samples (2 s each) emitted every ~150 ms = ~13× real-time synthesis. The run-3 total jumps because temperature=0.9 in the shipped sampling config happened to pick a longer trajectory — same prompt, different durations across runs.

Base + precomputed x-vector (voice cloning)

Model: Qwen/Qwen3-TTS-12Hz-0.6B-Base, five precomputed voices (one from Qwen's reference clip, four generated from OpenAI's TTS).

Voice Source Mode TTFA p50 TTFA min TTFA max Total p50
qwen_sample_m Qwen's clone_2.wav (~8 s English, male) icl 59 ms 54 66 404 ms
oai_echo OpenAI TTS, voice=echo (male) icl 61 ms 59 66 357 ms
oai_nova OpenAI TTS, voice=nova (female) icl 61 ms 58 63 421 ms
oai_onyx OpenAI TTS, voice=onyx (male) icl 62 ms 56 64 446 ms
oai_shimmer OpenAI TTS, voice=shimmer (female) icl 63 ms 55 68 409 ms
overall 61 ms 54 68 —

The ~10 ms gap vs. CustomVoice is Talker ref_code prefill — the ICL path carries 50-100 extra codec frames through the prompt. Switching the precompute to --mode xvec (speaker embedding only, no ref codes) closes the gap, with a modest fidelity cost on hard-to-clone voices.

The important part: precomputed voices get the same per-request cost as predefined speakers. All the heavy lifting (speaker encoder, codec encoder) happens once at precompute time (~15 s per voice on the 5090), results are cached to safetensors, and loaded into the engine's SpeakerEmbeddingCache at startup. Request time is a dict lookup.

What's left on the table

The vllm-omni team merged a qwen3_tts_fused_single_gpu.yaml deploy profile the day I was running this — single-stage, in-Talker streaming codec decoder, no shared-memory handoff. It's not in the 0.30.0 PyPI wheel yet but should drop next release.

The mrv2 pipeline I'm measuring here pays for a cross-process hop that the fused profile removes. Specifically, the shared-memory connector between the Talker and Code2Wav processes polls with connector_get_sleep_s: 0.01 (a 10 ms sleep per check) in the shipped qwen3_tts.yaml; add OS scheduling of the Code2Wav process on wake (~1-5 ms), D2H + serialization of the PCM payload back to the client-visible generator (~1-3 ms), and you're looking at ~5-15 ms of connector-hop overhead that the fused profile should recover. The upstream PR #8259 describes the architecture change in detail but doesn't publish benchmark numbers, so take this as a first-principles estimate, not a target. If you reproduce the setup after the fused profile ships in a vllm-omni release, you can swap DEPLOY_CONFIG_NAME in demo.py to qwen3_tts_fused_single_gpu.yaml and see where TTFA lands — somewhere in the 35-45 ms range on this hardware is the expectation, modulo whether the talker_stream_first_audio opt-in is enabled.

Voice cloning without the per-request cost

The naive Base-mode request looks like this:

{
"task_type": ["Base"],
"ref_audio": ["https://.../clone.wav"], # downloaded + codec-encoded per request
"ref_text": ["transcript of ref audio"],
"text": ["what you want synthesized"],
...
}

With a 5-second ref clip, the codec encode alone costs 50-80 ms on the 5090 — effectively doubling TTFA. For a serving system this is the wrong shape: the ref doesn't change between requests, so re-encoding it every time is just waste.

vLLM-Omni ships a precompute script (examples/online_serving/.../precompute_custom_voice.py) that does the one-time work: given a Base checkpoint + a ref audio clip, extract the 1024-dim speaker x-vector (and optionally the ref codec frames for ICL mode) and save them as safetensors. Point custom_voice_dir at the directory of precomputed voices in your deploy YAML, and vLLM-Omni will preload everything into its SpeakerEmbeddingCache at engine startup. Requests reference voices by name instead of by ref_audio.

My pipeline:

reference_voices/oai_echo.wav + oai_echo.txt
          │  (one-time, ~15 s per voice on 5090)
          ▼
precompute_voice.py  ─►  voices/oai_echo.safetensors
                         voices/custom_voice_manifest.json
          │  (engine startup — ~5 min for CUDA graphs)
          ▼
AsyncOmni(deploy_config="qwen3_tts_base_voices.yaml")
          │  (per request — O(1) dict lookup, no codec encoder)
          ▼
TTFA ≈ 59-62 ms

For the reference clips themselves I used OpenAI's TTS to synthesize four voices (two male, two female). My API key stays in OPENAI_API_KEY, never touches disk, logs, or any committed artifact. The resulting WAVs are OpenAI output — safe to commit, which is how this project ships them for anyone cloning the repo to listen to without having to re-run the OpenAI step.

Making it audible — a side journey through audio stacks

The TTFA numbers above describe when the first PCM chunk hits your Python iterator, not when the sound actually reaches a human ear. If you build a REPL that just does sd.OutputStream.write(chunk) as chunks arrive, you learn — the hard way, over four debugging passes — that getting from "chunk arrives" to "clean sound" has its own gauntlet. Writing this up because the fixes are not well-documented anywhere I could find.

1. The first chunk is 80 ms but the second doesn't arrive for 150 ms

Qwen's streaming profile emits a tiny first chunk (initial_codec_chunk_frames: 1) for low TTFA — 1920 samples of audio at 24 kHz, so 80 ms. The second chunk is 48,000 samples (2 s) and arrives ~150 ms wall-clock later. If you start playback the instant the first chunk arrives with latency="low" (PortAudio's 5-20 ms ring), the DAC drains the 80 ms chunk in 80 ms and then runs dry for ~70 ms before chunk 2 is written. Audible click. The naive fix (bigger buffer, latency=0.2) just moves the problem — the DAC starts draining immediately and the ring runs dry mid-utterance anyway.

Real fix: pre-buffer in Python before calling stream.start(). Collect chunks until you've got ~250 ms queued, then start the stream and write everything at once. Reported TTFA (first-chunk arrival from the model) is unchanged; perceived audio starts ~200 ms later. No clicks.

2. The first phoneme of the first word gets eaten

Even with the pre-buffer fix, the first prompt of a fresh REPL session loses its first ~50 ms of speech. Diagnosis: OS audio stacks (ALSA → PipeWire/PulseAudio → DAC) take 30-80 ms to physically produce sound from a cold start. During that time PortAudio happily drains buffer samples into the void. Qwen's output has no lead-in silence — speech starts at sample 0 — so the first phonemes vanish.

Partial fix: prepend 100 ms of silence to the first write, so the hardware warms up on silence instead of on real audio.

3. Really fix it: keep the stream open

The silence pre-roll works, but it pays the cold-start on every utterance (because the stream closes between prompts). The right design is a persistent sd.OutputStream with a callback that lives for the whole REPL session:

class StreamingPlayer:
def __init__(self, ...):
self._queue = queue.Queue()
self._pending = np.zeros(0, dtype=np.float32) # callback-only state
...

def _callback(self, outdata, frames, time_info, status):
outdata.fill(0) # default to silence when queue is dry
written = 0
if self._pending.size > 0:
take = min(self._pending.size, frames)
outdata[:take, 0] = self._pending[:take]
self._pending = self._pending[take:]
written = take
while written < frames:
try:
chunk = self._queue.get_nowait()
except queue.Empty:
break
take = min(chunk.size, frames - written)
outdata[written:written+take, 0] = chunk[:take]
if take < chunk.size:
self._pending = chunk[take:]
written += take

Opened once at REPL startup, closed at exit. Between prompts the callback fills with silence — so the DAC stays warm. By the time the user types their first prompt (which happens after the ~5 min CUDA-graph warmup), the audio chain has been producing sound for minutes. First phoneme of every prompt lands intact. Pre-buffer disappears; silence pre-roll disappears; code simplifies. This is the design you want for any real application.

4. The metallic artifact

With a working REPL you still occasionally hear a shimmering/metallic quality on longer utterances, getting worse as speech progresses. This isn't a chunk-boundary click or a model artifact. It's resampling aliasing. Your consumer audio device almost certainly has a native rate of 48 kHz or 44.1 kHz, not 24 kHz. If you open an sd.OutputStream(samplerate=24000) on a 44.1 kHz device, the OS audio stack has to resample 24000 → 44100 using whatever filter it has lying around — and ALSA's default is linear interpolation. The ratio 24/44.1 is not clean, so the filter accumulates phase error over long passes, which an ear perceives as a progressive metallic ring.

Fix: resample in Python with a high-quality filter and open the stream at an OS-native rate. I use soxr.ResampleStream(quality="VHQ") to go 24 kHz → 48 kHz in-process. 48 kHz is universally supported, 24 → 48 is a clean 2× ratio, and VHQ polyphase is effectively indistinguishable from no resample at all. The OS passes it through:

Qwen 24 kHz → soxr VHQ 2× upsample (Python) → sd.OutputStream @ 48 kHz
            → ALSA/PipeWire passthrough → DAC

Metallic artifact disappears. For completeness, you also need to call drain_resampler() at the end of each utterance so the polyphase filter's internal history (~a dozen samples) flushes to the DAC — otherwise the trailing consonant of each utterance gets swallowed by group delay.

Did the audio fixes impact the benchmark?

No. All of the above lives in the playback path. The benchmark scripts (demo.py, clone_bench.py) capture chunks directly into a timing structure and never touch sd.OutputStream. Rerun of the full bench after the fixes:

TTFA p50 before TTFA p50 after
CustomVoice (Ryan) 47 ms 51 ms
Base clone (qwen_sample_m) 59 ms 59 ms

The ~4 ms CustomVoice drift is noise plus mild GPU-0 contention from an unrelated workload sharing the host; p50 bounces in the 47-51 ms band across runs regardless of fixes. Base clone is flat to within sampling noise. No regression on headline latency numbers from any of the playback work.

Subjective impressions

To be honest, I'm not wild about the voices included in the Customvoice model. One issue is that they only include two native-english speaking voices, they're both male and they aren't to my taste. The other voices produce a marked accent when reading English.

I got much better results with the voice cloning model. I generated the voice-cloning reference clips using OpenAI's TTS model and generated two male and two female voices. Using these reference clips along with the model to generate speech sounded great to my ears. I'll be pushing on using this model in additional applications given these results.

Honest limitations

Reproduce it

Everything is in the repo. Clone, follow the README.md quick start, and you should see numbers within a few ms of what's reported here on comparable Blackwell hardware. Non-Blackwell GPUs work too — change TORCH_CUDA_ARCH_LIST and the shipped mrv2 config will do the right thing.

Repo: https://github.com/johnrobinsn/qwen3-tts-stream-demo

Samples live under showcase/:

All files are 24 kHz 16-bit PCM WAV, playable in any browser or player.


Get more content like this. Occasional, low-cadence, real-news-only.


Credits

Qwen team for the model and the dual-track streaming architecture (paper arXiv:2601.15621). vLLM-Omni team for the production inference stack with day-0 Qwen3-TTS support (github.com/vllm-project/vllm-omni). OpenAI for the TTS used to generate reference clips for cloning.


[1] https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice — model card. The claim appears verbatim in the "Low Latency" bullet.


Share on Twitter |  Discuss on Twitter

John Robinson © 2022-2026