I-Parakeet is an integer-only version of Nvidia’s 600-million-parameter Parakeet-CTC speech-recognition model that runs on a smartphone neural processing unit without sending sensitive operations back to a floating-point CPU. Eliminating those fallbacks allows the model to use the mobile accelerator throughout inference.
The researchers reformulated the Conformer model’s relative-position attention as integer operations, including branches that use different quantization scales. They also created an integer approximation for the Swish activation function and used targeted 16-bit handling and percentile calibration where activation ranges made lower precision unreliable.
On LibriSpeech’s difficult test-other split, I-Parakeet recorded a 4.97% word error rate. It ran on a Qualcomm NPU with a real-time factor of 0.048—about 7.5 times faster than the paper’s CPU baseline, meaning it processed audio substantially faster than playback speed. Results from one benchmark and device do not establish performance across accents, languages, or every phone. The work is notable because it preserves a large recognizer’s architecture while making every operator compatible with integer-oriented mobile hardware, reducing dependence on mixed execution paths that undermine edge efficiency.