Multimodal LLMs vs Dedicated ASR for Russian Speech: 9 Engines on One 16GB GPU

Multimodal LLMs vs Dedicated ASR for Russian Speech: 9 Engines on One 16GB GPU

Every audio-capable LLM release comes with a claim that it "understands speech". That is not the same claim as "it can replace your ASR service". The difference only shows up when you put both on the same card, feed them the same audio, and compute the same metric.

This article does that on one RTX 5060 Ti (16 GB), with two Russian corpora, nine engines, and per-sample transcripts kept for recomputation. The short version:

  • Dedicated ASR wins the thing you deploy: GigaAM v3 transcribes at 34–40× real time in 674 MiB of VRAM. The best multimodal LLM needs ~9.2 GiB and runs at 6.7–8×.
  • On clean studio reads the gap is not proven. At n=50 the paired interval between Gemma 4 12B QAT and GigaAM is +0.28 п.п. [-1.60; +2.46] — statistically indistinguishable.
  • On spontaneous speech it is proven and large. Every multimodal model is significantly worse; Qwen2.5-Omni is +16.2 п.п. behind GigaAM.
  • Quantization mode decides whether the model transcribes at all. The same Gemma 4 E4B checkpoint scores 7.3% WER in bnb 4-bit and 75.6% in bnb 8-bit. Bit width is not the axis that matters.

The setup

One GPU, one engine at a time. Everything else about the host is held constant.

ParameterValue
GPURTX 5060 Ti, 16 GB (GPU1 of a 2-card host; the 4090 serves LLMs and was untouched)
ServingDocker, runtime: nvidia, NVIDIA_VISIBLE_DEVICES=1, one OpenAI-compatible ASR container per port
ProtocolPOST /v1/audio/transcriptions, multipart WAV, language=ru, response_format=json
IsolationSpeaches/GigaAM stopped while an LLM occupies the card; VRAM verified free between switches
Clientharness/run_bench.py — one process, sequential requests, no concurrency
Scoringharness/analyze.py — jiwer WER/CER, bootstrap CI over utterances (20,000 resamples), paired per-utterance diffs

The engines and how they were loaded:

EnginePortLoading
speaches8002deepdml/faster-whisper-large-v3-turbo-ct2 (CTranslate2)
gigaam-local8003GigaAM v3 v3_e2e_rnnt (SberDevices, MIT)
gemma4-e2b8004google/gemma-4-E2B-it, bf16
gemma4-e4b-4bit8004google/gemma-4-E4B-it, bitsandbytes NF4
gemma4-e4b-8bit8006same checkpoint, bitsandbytes int8
gemma4-e4b-qat8004Google QAT release + bnb4
gemma4-12b-qat11434Ollama gemma4:12b-it-qat (official QAT Q4_0 GGUF)
qwen2-audio-7b8005Qwen/Qwen2-Audio-7B-Instruct, bnb4
qwen25-omni-7b8007Qwen/Qwen2.5-Omni-7B, bnb4 (see the GPTQ section below)

Corpora

Two domains, deliberately not averaged together:

  • FLEURS-ru (google/fleurs, ru_ru, test) — studio reading of news sentences. n=50, 621.2 s of audio.
  • Podlodka Speech (bond005/podlodka_speech, test) — podcast lectures and interviews, spontaneous speech. n=20, 486.5 s of audio.

A single corpus would have produced a cleaner story and a wrong one. Clean reads and spontaneous speech are different tasks, and as shown below the ranking gap between them is the actual result.

Normalization

Two normalizations are computed for every engine:

  1. Base — lowercase, ё→е, punctuation stripped, whitespace collapsed.
  2. +numbers — the same, plus spelled-out numerals converted to digits.

The second column exists because engines disagree on formatting, not on hearing. FLEURS references write numbers as digits, so a model that answers in words is penalized for style. Qwen2.5-Omni is the clearest case in this run: 7.19% base → 5.77% with number parsing, a 1.4-point improvement that has nothing to do with acoustics. For Russian the conversion is incomplete (cardinals parse, ordinals do not), so the numbers column supplements the base column rather than replacing it.

Results: FLEURS n=50 (studio reads)

EngineWER95% CI+numbers×RTmed latVRAM
speaches4.07%[2.62; 5.77]4.07%12.9×0.92s~2208 MiB
gigaam-local5.11%[3.29; 7.25]5.11%34.1×0.36s~674 MiB
gemma4-12b-qat5.39%[3.27; 8.08]5.39%6.7×1.60s~9193 MiB
gemma4-e2b5.96%[3.99; 8.25]5.96%12.4×0.94s~10129 MiB
gemma4-e4b-qat7.10%[4.68; 9.97]7.10%6.8×1.68s~9365 MiB
qwen25-omni-7b7.19%[4.85; 9.78]5.77%11.3×1.06s~7000–8340 MiB
gemma4-e4b-4bit7.28%[5.10; 9.90]7.10%7.1×1.61s~9261 MiB
qwen2-audio-7b31.50%[25.82; 38.04]31.32%5.3×2.15s~6615–7777 MiB
gemma4-e4b-8bit75.59%[65.58; 85.05]75.59%5.4×1.66s~11269 MiB

At n=50 the top of this table is a cluster, not a ranking. Paired intervals on the same utterances:

text
gemma4-12b-qat − gigaam-local     +0.28 п.п. [-1.60; +2.46]  not proven
gemma4-12b-qat − speaches         +1.32 п.п. [-0.67; +3.77]  not proven
gemma4-12b-qat − qwen25-omni-7b   -1.80 п.п. [-4.10; +0.68]  not proven
gemma4-e4b-qat − speaches         +3.03 п.п. [+1.07; +5.21]  significant
gigaam-local   − qwen25-omni-7b   -2.08 п.п. [-3.76; -0.48]  significant

So the honest claim is: on clean read speech, a QAT 12B multimodal model is not measurably worse than a dedicated ASR — while costing 14× the VRAM and 5× the latency.

Results: Podlodka n=20 (spontaneous speech)

EngineWER95% CI×RTmed lat
gigaam-local7.30%[4.69; 10.01]40.5×0.51s
speaches7.38%[5.04; 10.03]9.1×1.66s
gemma4-12b-qat10.73%[7.69; 14.88]8.0×2.85s
gemma4-e4b-4bit11.07%[8.50; 14.06]6.3×3.46s
gemma4-e4b-qat11.67%[8.60; 14.88]6.3×3.50s
qwen25-omni-7b23.52%[13.25; 33.89]13.5×1.67s
qwen2-audio-7b35.88%[29.83; 41.32]5.3×4.30s
gemma4-e4b-8bit63.43%[52.94; 72.65]4.6×5.21s

Here the paired tests stop hedging:

text
gemma4-12b-qat − gigaam-local     +3.43 п.п. [+0.74; +6.77]   significant
gemma4-e4b-qat − gigaam-local     +4.38 п.п. [+2.83; +5.90]   significant
qwen25-omni-7b − speaches        +16.14 п.п. [+6.41; +26.19]  significant
gigaam-local   − speaches         -0.09 п.п. [-1.67; +1.36]   not proven

Qwen2.5-Omni loses 16 points to a 674 MiB model the moment speech stops being dictated. Its FLEURS score (7.19%) and its Podlodka score (23.52%) come from the same weights and the same prompt; only the speech register changed. That is the whole argument for per-domain benchmarking in one pair of numbers.

What broke along the way

The results above took more debugging than measuring. Each of these is a real failure mode, not a configuration typo.

Qwen2.5-Omni GPTQ-Int4 does not load with from_pretrained

The official Qwen/Qwen2.5-Omni-7B-GPTQ-Int4 card promises 11.6 GiB for a 15-second clip, which is exactly what a 16 GB card wants. It does not load through plain transformers: the checkpoint expects the repo's low-VRAM path (GPTQModel.load with a hand-written device_map, a monkeypatched from_config that also loads spk_dict.pt, and a streaming token2wav). Falling back to Qwen/Qwen2.5-Omni-7B with bitsandbytes NF4 gave a working server at ~7–8.3 GiB, which is what the numbers above measure. The GPTQ path is worth revisiting, but it is a port, not a flag.

disable_talker() breaks generate()

The documented way to save ~2 GB when you only need text is model.disable_talker(). With the current transformers, generation then dies:

text
AttributeError: 'Qwen2_5OmniForConditionalGeneration' object has no attribute 'talker'
  at modeling_qwen2_5_omni.py: "pad_token_id": self.talker.codec_pad_token

generate() reads self.talker.codec_pad_token even with return_audio=False. The workaround is to keep the talker loaded (pay the ~700 MiB) and simply never request audio.

Chat templates leak into transcripts

The first working Omni response was:

text
"system You are a speech recognition model. user Transcribe the Russian audio ... assistant вот"

batch_decode returns the whole sequence, prompt included. Any LLM-as-ASR wrapper has to slice by prompt length before decoding, or the WER number measures your template:

python
prompt_len = int(inputs["input_ids"].shape[-1])
if text_ids.shape[-1] > prompt_len:
    text_ids = text_ids[:, prompt_len:]

Ollama: use the chat API, not the audio endpoint

For gemma4:12b-it-qat, POST /v1/audio/transcriptions hung. The working path is the chat API with an input_audio content part, thinking disabled (reasoning_effort: none) and a bounded num_ctx (8192). Without disabling reasoning, the model spends its budget thinking about the audio instead of transcribing it.

8-bit is not "less lossy than 4-bit"

The single most counter-intuitive result of the run:

Gemma 4 E4BPodlodka WERFLEURS WER
bnb 4-bit (NF4)11.07%7.28%
bnb 8-bit (int8)63.43%75.59%

The 8-bit build even uses more VRAM (11.3 GiB vs 9.3 GiB) and is slower. Both domains agree, and the paired interval is +68.5 п.п. [+59.0; +77.5] — this is not sampling noise, it is a broken deployment mode for this architecture. The corollary: never assume a quantization is safe because a higher-precision label sounds safer. Measure the mode you actually ship.

QAT is not a synonym for "4-bit"

Same lesson from the other direction. Gemma 4 12B in generic bnb NF4 scores 34.15% / 45.92% — unusable. The official QAT Q4_0 GGUF of the same size class scores 5.39% / 10.73% and became the best multimodal engine in this benchmark. Quantization-aware training and post-training quantization are not interchangeable for audio.

Qwen2-Audio silently ignores audio

Qwen2AudioProcessor takes audio=, not audios=. Passing the wrong keyword produces no error, no warning, and fluent Russian text that has nothing to do with the recording. Additionally, the base checkpoint echoes instructions back; only -Instruct behaves as a transcriber. A "model that hallucinates" is sometimes a processor that never received the waveform.

Dataset delivery is part of the harness

Streaming FLEURS from the Hub timed out repeatedly mid-run (The read operation timed out on a 539 MB parquet), producing a zero-row result file that looked like an engine failure. Fixing it meant downloading the parquet with a resumable client, placing it in the Hub cache layout, and reading the split locally. If a run can be invalidated by someone else's CDN, that is a benchmark bug, not bad luck.

What this means in practice

  • For production Russian ASR, use dedicated ASR. GigaAM v3 is the whole argument: best or tied-best on both domains, 40× real time, 674 MiB. On a 16 GB card it leaves room for everything else you want to run. On timestamps: the official GigaAM API can return word-level timings (transcribe(..., word_timestamps=True)) and, for long audio, VAD-based segment bounds via transcribe_longform (needs pyannote). That landed in the upstream package around 2026/04. Whether you see them over HTTP depends on the wrapper: a plain response_format=json OpenAI call (as in this harness) still returns a flat string; servers that implement verbose_json / timestamp_granularities can expose segments and words. For call timelines you still need speaker diarization on top — ASR timings alone do not label who spoke.
  • Multimodal LLMs are worth it when you need more than a transcript. If the same call must transcribe and summarize, classify or extract structure, gemma4-12b-qat at ~9.2 GiB is the first option in this set whose clean-speech accuracy is not measurably worse than an ASR service. You still pay ~5× latency and a 3.4-point penalty on spontaneous speech.
  • Do not trust a single-corpus claim, including your own. Omni moves from "competitive" (7.19%) to "unusable" (23.52%) between two corpora with no configuration change.
  • Co-residency is the real constraint. GigaAM (674 MiB) plus Speaches (2.2 GiB) fit together with 13 GiB to spare. Any of the Gemma builds needs the card almost to itself, so "just add multimodal alongside" is a VRAM decision before it is an accuracy decision.

Lessons

  • Paired confidence intervals change conclusions. At n=50 the FLEURS top-four are one cluster; reporting the ordering as a ranking would be wrong. At n=20 on spontaneous speech the multimodal gap is significant anyway — small n does not automatically mean "no result", it means check.
  • Report two normalizations. A 1.4-point swing from number formatting (Omni: 7.19% → 5.77%) is the difference between two adjacent ranks, and it measures style, not recognition.
  • Wrapper bugs look exactly like model weakness. Prompt leakage, a wrong processor keyword and a hung endpoint each produced plausible-looking bad WER before they produced an error message.
  • Keep per-sample transcripts. Every number here is recomputable from JSONL, which is what made the 8-bit result believable instead of suspicious.

Resources

Published on 9/17/2026

Loading tags...