
Multimodal LLMs vs Dedicated ASR for Russian Speech: 9 Engines on One 16GB GPU
Every audio-capable LLM release comes with a claim that it "understands speech". That is not the same claim as "it can replace your ASR service". The difference only shows up when you put both on the same card, feed them the same audio, and compute the same metric.
This article does that on one RTX 5060 Ti (16 GB), with two Russian corpora, nine engines, and per-sample transcripts kept for recomputation. The short version:
- Dedicated ASR wins the thing you deploy: GigaAM v3 transcribes at 34–40× real time in 674 MiB of VRAM. The best multimodal LLM needs ~9.2 GiB and runs at 6.7–8×.
- On clean studio reads the gap is not proven. At n=50 the paired
interval between Gemma 4 12B QAT and GigaAM is
+0.28 п.п. [-1.60; +2.46]— statistically indistinguishable. - On spontaneous speech it is proven and large. Every multimodal model is significantly worse; Qwen2.5-Omni is +16.2 п.п. behind GigaAM.
- Quantization mode decides whether the model transcribes at all. The same Gemma 4 E4B checkpoint scores 7.3% WER in bnb 4-bit and 75.6% in bnb 8-bit. Bit width is not the axis that matters.
The setup
One GPU, one engine at a time. Everything else about the host is held constant.
| Parameter | Value |
|---|---|
| GPU | RTX 5060 Ti, 16 GB (GPU1 of a 2-card host; the 4090 serves LLMs and was untouched) |
| Serving | Docker, runtime: nvidia, NVIDIA_VISIBLE_DEVICES=1, one OpenAI-compatible ASR container per port |
| Protocol | POST /v1/audio/transcriptions, multipart WAV, language=ru, response_format=json |
| Isolation | Speaches/GigaAM stopped while an LLM occupies the card; VRAM verified free between switches |
| Client | harness/run_bench.py — one process, sequential requests, no concurrency |
| Scoring | harness/analyze.py — jiwer WER/CER, bootstrap CI over utterances (20,000 resamples), paired per-utterance diffs |
The engines and how they were loaded:
| Engine | Port | Loading |
|---|---|---|
speaches | 8002 | deepdml/faster-whisper-large-v3-turbo-ct2 (CTranslate2) |
gigaam-local | 8003 | GigaAM v3 v3_e2e_rnnt (SberDevices, MIT) |
gemma4-e2b | 8004 | google/gemma-4-E2B-it, bf16 |
gemma4-e4b-4bit | 8004 | google/gemma-4-E4B-it, bitsandbytes NF4 |
gemma4-e4b-8bit | 8006 | same checkpoint, bitsandbytes int8 |
gemma4-e4b-qat | 8004 | Google QAT release + bnb4 |
gemma4-12b-qat | 11434 | Ollama gemma4:12b-it-qat (official QAT Q4_0 GGUF) |
qwen2-audio-7b | 8005 | Qwen/Qwen2-Audio-7B-Instruct, bnb4 |
qwen25-omni-7b | 8007 | Qwen/Qwen2.5-Omni-7B, bnb4 (see the GPTQ section below) |
Corpora
Two domains, deliberately not averaged together:
- FLEURS-ru (
google/fleurs,ru_ru, test) — studio reading of news sentences. n=50, 621.2 s of audio. - Podlodka Speech (
bond005/podlodka_speech, test) — podcast lectures and interviews, spontaneous speech. n=20, 486.5 s of audio.
A single corpus would have produced a cleaner story and a wrong one. Clean reads and spontaneous speech are different tasks, and as shown below the ranking gap between them is the actual result.
Normalization
Two normalizations are computed for every engine:
- Base — lowercase,
ё→е, punctuation stripped, whitespace collapsed. - +numbers — the same, plus spelled-out numerals converted to digits.
The second column exists because engines disagree on formatting, not on hearing. FLEURS references write numbers as digits, so a model that answers in words is penalized for style. Qwen2.5-Omni is the clearest case in this run: 7.19% base → 5.77% with number parsing, a 1.4-point improvement that has nothing to do with acoustics. For Russian the conversion is incomplete (cardinals parse, ordinals do not), so the numbers column supplements the base column rather than replacing it.
Results: FLEURS n=50 (studio reads)
| Engine | WER | 95% CI | +numbers | ×RT | med lat | VRAM |
|---|---|---|---|---|---|---|
| speaches | 4.07% | [2.62; 5.77] | 4.07% | 12.9× | 0.92s | ~2208 MiB |
| gigaam-local | 5.11% | [3.29; 7.25] | 5.11% | 34.1× | 0.36s | ~674 MiB |
| gemma4-12b-qat | 5.39% | [3.27; 8.08] | 5.39% | 6.7× | 1.60s | ~9193 MiB |
| gemma4-e2b | 5.96% | [3.99; 8.25] | 5.96% | 12.4× | 0.94s | ~10129 MiB |
| gemma4-e4b-qat | 7.10% | [4.68; 9.97] | 7.10% | 6.8× | 1.68s | ~9365 MiB |
| qwen25-omni-7b | 7.19% | [4.85; 9.78] | 5.77% | 11.3× | 1.06s | ~7000–8340 MiB |
| gemma4-e4b-4bit | 7.28% | [5.10; 9.90] | 7.10% | 7.1× | 1.61s | ~9261 MiB |
| qwen2-audio-7b | 31.50% | [25.82; 38.04] | 31.32% | 5.3× | 2.15s | ~6615–7777 MiB |
| gemma4-e4b-8bit | 75.59% | [65.58; 85.05] | 75.59% | 5.4× | 1.66s | ~11269 MiB |
At n=50 the top of this table is a cluster, not a ranking. Paired intervals on the same utterances:
gemma4-12b-qat − gigaam-local +0.28 п.п. [-1.60; +2.46] not proven
gemma4-12b-qat − speaches +1.32 п.п. [-0.67; +3.77] not proven
gemma4-12b-qat − qwen25-omni-7b -1.80 п.п. [-4.10; +0.68] not proven
gemma4-e4b-qat − speaches +3.03 п.п. [+1.07; +5.21] significant
gigaam-local − qwen25-omni-7b -2.08 п.п. [-3.76; -0.48] significant
So the honest claim is: on clean read speech, a QAT 12B multimodal model is not measurably worse than a dedicated ASR — while costing 14× the VRAM and 5× the latency.
Results: Podlodka n=20 (spontaneous speech)
| Engine | WER | 95% CI | ×RT | med lat |
|---|---|---|---|---|
| gigaam-local | 7.30% | [4.69; 10.01] | 40.5× | 0.51s |
| speaches | 7.38% | [5.04; 10.03] | 9.1× | 1.66s |
| gemma4-12b-qat | 10.73% | [7.69; 14.88] | 8.0× | 2.85s |
| gemma4-e4b-4bit | 11.07% | [8.50; 14.06] | 6.3× | 3.46s |
| gemma4-e4b-qat | 11.67% | [8.60; 14.88] | 6.3× | 3.50s |
| qwen25-omni-7b | 23.52% | [13.25; 33.89] | 13.5× | 1.67s |
| qwen2-audio-7b | 35.88% | [29.83; 41.32] | 5.3× | 4.30s |
| gemma4-e4b-8bit | 63.43% | [52.94; 72.65] | 4.6× | 5.21s |
Here the paired tests stop hedging:
gemma4-12b-qat − gigaam-local +3.43 п.п. [+0.74; +6.77] significant
gemma4-e4b-qat − gigaam-local +4.38 п.п. [+2.83; +5.90] significant
qwen25-omni-7b − speaches +16.14 п.п. [+6.41; +26.19] significant
gigaam-local − speaches -0.09 п.п. [-1.67; +1.36] not proven
Qwen2.5-Omni loses 16 points to a 674 MiB model the moment speech stops being dictated. Its FLEURS score (7.19%) and its Podlodka score (23.52%) come from the same weights and the same prompt; only the speech register changed. That is the whole argument for per-domain benchmarking in one pair of numbers.
What broke along the way
The results above took more debugging than measuring. Each of these is a real failure mode, not a configuration typo.
Qwen2.5-Omni GPTQ-Int4 does not load with from_pretrained
The official Qwen/Qwen2.5-Omni-7B-GPTQ-Int4 card promises 11.6 GiB for a
15-second clip, which is exactly what a 16 GB card wants. It does not load
through plain transformers: the checkpoint expects the repo's low-VRAM path
(GPTQModel.load with a hand-written device_map, a monkeypatched
from_config that also loads spk_dict.pt, and a streaming token2wav).
Falling back to Qwen/Qwen2.5-Omni-7B with bitsandbytes NF4 gave a working
server at ~7–8.3 GiB, which is what the numbers above measure. The GPTQ path is
worth revisiting, but it is a port, not a flag.
disable_talker() breaks generate()
The documented way to save ~2 GB when you only need text is
model.disable_talker(). With the current transformers, generation then
dies:
AttributeError: 'Qwen2_5OmniForConditionalGeneration' object has no attribute 'talker'
at modeling_qwen2_5_omni.py: "pad_token_id": self.talker.codec_pad_token
generate() reads self.talker.codec_pad_token even with
return_audio=False. The workaround is to keep the talker loaded (pay the
~700 MiB) and simply never request audio.
Chat templates leak into transcripts
The first working Omni response was:
"system You are a speech recognition model. user Transcribe the Russian audio ... assistant вот"
batch_decode returns the whole sequence, prompt included. Any LLM-as-ASR
wrapper has to slice by prompt length before decoding, or the WER number
measures your template:
prompt_len = int(inputs["input_ids"].shape[-1])
if text_ids.shape[-1] > prompt_len:
text_ids = text_ids[:, prompt_len:]
Ollama: use the chat API, not the audio endpoint
For gemma4:12b-it-qat, POST /v1/audio/transcriptions hung. The working path
is the chat API with an input_audio content part, thinking disabled
(reasoning_effort: none) and a bounded num_ctx (8192). Without disabling
reasoning, the model spends its budget thinking about the audio instead of
transcribing it.
8-bit is not "less lossy than 4-bit"
The single most counter-intuitive result of the run:
| Gemma 4 E4B | Podlodka WER | FLEURS WER |
|---|---|---|
| bnb 4-bit (NF4) | 11.07% | 7.28% |
| bnb 8-bit (int8) | 63.43% | 75.59% |
The 8-bit build even uses more VRAM (11.3 GiB vs 9.3 GiB) and is slower.
Both domains agree, and the paired interval is +68.5 п.п. [+59.0; +77.5] —
this is not sampling noise, it is a broken deployment mode for this
architecture. The corollary: never assume a quantization is safe because a
higher-precision label sounds safer. Measure the mode you actually ship.
QAT is not a synonym for "4-bit"
Same lesson from the other direction. Gemma 4 12B in generic bnb NF4 scores 34.15% / 45.92% — unusable. The official QAT Q4_0 GGUF of the same size class scores 5.39% / 10.73% and became the best multimodal engine in this benchmark. Quantization-aware training and post-training quantization are not interchangeable for audio.
Qwen2-Audio silently ignores audio
Qwen2AudioProcessor takes audio=, not audios=. Passing the wrong keyword
produces no error, no warning, and fluent Russian text that has nothing to do
with the recording. Additionally, the base checkpoint echoes instructions back;
only -Instruct behaves as a transcriber. A "model that hallucinates" is
sometimes a processor that never received the waveform.
Dataset delivery is part of the harness
Streaming FLEURS from the Hub timed out repeatedly mid-run
(The read operation timed out on a 539 MB parquet), producing a zero-row
result file that looked like an engine failure. Fixing it meant downloading the
parquet with a resumable client, placing it in the Hub cache layout, and
reading the split locally. If a run can be invalidated by someone else's CDN,
that is a benchmark bug, not bad luck.
What this means in practice
- For production Russian ASR, use dedicated ASR. GigaAM v3 is the whole
argument: best or tied-best on both domains, 40× real time, 674 MiB. On a
16 GB card it leaves room for everything else you want to run.
On timestamps: the official GigaAM API can return word-level timings
(
transcribe(..., word_timestamps=True)) and, for long audio, VAD-based segment bounds viatranscribe_longform(needs pyannote). That landed in the upstream package around 2026/04. Whether you see them over HTTP depends on the wrapper: a plainresponse_format=jsonOpenAI call (as in this harness) still returns a flat string; servers that implementverbose_json/timestamp_granularitiescan expose segments and words. For call timelines you still need speaker diarization on top — ASR timings alone do not label who spoke. - Multimodal LLMs are worth it when you need more than a transcript.
If the same call must transcribe and summarize, classify or extract
structure,
gemma4-12b-qatat ~9.2 GiB is the first option in this set whose clean-speech accuracy is not measurably worse than an ASR service. You still pay ~5× latency and a 3.4-point penalty on spontaneous speech. - Do not trust a single-corpus claim, including your own. Omni moves from "competitive" (7.19%) to "unusable" (23.52%) between two corpora with no configuration change.
- Co-residency is the real constraint. GigaAM (674 MiB) plus Speaches (2.2 GiB) fit together with 13 GiB to spare. Any of the Gemma builds needs the card almost to itself, so "just add multimodal alongside" is a VRAM decision before it is an accuracy decision.
Lessons
- Paired confidence intervals change conclusions. At n=50 the FLEURS top-four are one cluster; reporting the ordering as a ranking would be wrong. At n=20 on spontaneous speech the multimodal gap is significant anyway — small n does not automatically mean "no result", it means check.
- Report two normalizations. A 1.4-point swing from number formatting (Omni: 7.19% → 5.77%) is the difference between two adjacent ranks, and it measures style, not recognition.
- Wrapper bugs look exactly like model weakness. Prompt leakage, a wrong processor keyword and a hung endpoint each produced plausible-looking bad WER before they produced an error message.
- Keep per-sample transcripts. Every number here is recomputable from JSONL, which is what made the 8-bit result believable instead of suspicious.
Resources
- stt-ru-benchmark (harness and per-sample results): https://github.com/berdachuk/stt-ru-benchmark
- GigaAM (SberDevices): https://github.com/salute-developers/GigaAM
- faster-whisper large-v3-turbo CT2: https://huggingface.co/deepdml/faster-whisper-large-v3-turbo-ct2
- Qwen2.5-Omni-7B: https://huggingface.co/Qwen/Qwen2.5-Omni-7B
- Qwen2.5-Omni-7B-GPTQ-Int4 (low-VRAM path): https://huggingface.co/Qwen/Qwen2.5-Omni-7B-GPTQ-Int4
- FLEURS: https://huggingface.co/datasets/google/fleurs
- Podlodka Speech: https://huggingface.co/datasets/bond005/podlodka_speech
- Bootstrap CI method (Bisani & Ney, ICASSP 2004): https://ieeexplore.ieee.org/document/1326009
Published on 9/17/2026