How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Faster Gemma 4 on MLX with multi-token prediction

calendar_today June 29, 2026 person domain ollama

Gemma 4 is now significantly faster in Ollama 0.31 on Apple Silicon via multi-token prediction (MTP), powered by MLX. Performance is up to 90% faster when used with coding agents, as measured on the Aider polyglot benchmark.

open_in_new Read original post