1 of 1
Story summary
- Google launched the Gemma 4 open-weight models with Multi-Token Prediction (MTP) drafters, promising up to a three-fold inference speedup.
- Google re-licensed Gemma 4 under the Apache 2.0 license and made it runnable on TPU chips or quantized for consumer GPUs.
- The MTP drafters use speculative decoding that halves token-generation wait time and can achieve up to a three-fold speedup.
- Engineers Maarten Grootendorst and Olivier Lacombe say the speedup causes no loss in output quality.
