Google Research described how frozen Multi-Token Prediction can accelerate Gemini Nano models on Pixel devices. The technique is aimed at improving local model speed while preserving the practical constraints of on-device deployment.

On-device AI matters because latency, privacy, and offline availability are difficult to solve if every interaction must round-trip to the cloud. Better local performance can make small models more useful in everyday mobile features.

The research also shows how model optimization is becoming device-specific: frontier cloud models get attention, but efficient local models determine what AI can run directly in users’ hands.