PrismML has compressed a 27-billion-parameter reasoning model to under 4GB, small enough to run on an iPhone, according to The Decoder. The company says the smallest Bonsai version preserves about 90% of the original model's performance in its own benchmarks.
The result points to a major direction for on-device AI: making stronger reasoning models small enough for consumer hardware. If the performance claims hold up, compressed open models could reduce dependence on cloud inference for some coding, math, and assistant tasks.