Articles tagged "speculative-decoding"

10-Jun-2026
Running Gemma 4 12B Locally with Speculative Decoding on Ubuntu
A practical guide to running Gemma 4 12B locally on Ubuntu using llama.cpp, Hugging Face automatic model loading, and speculative decoding with MTP.

15-May-2026
Running Gemma 4 31B Faster at Home with Speculative Decoding
A practical guide to running Gemma 4 31B on a dual-GPU Ubuntu box using llama.cpp TurboQuant, Gemma 4’s MTP assistant head, and speculative decoding to nearly double token throughput.