TechCrunch published an in-depth analysis of the growing cost crisis in AI inference, where the per-token pricing model is colliding with the reality of serving ever-larger models to millions of users. The report details how companies are scrambling to manage expenses that can spiral unpredictably as usage scales.

The article traces the economics from GPU procurement through model serving, revealing that even well-funded AI labs are finding inference costs to be their fastest-growing expense line. Techniques like quantization, speculative decoding, and caching are being deployed aggressively, but the fundamental cost-per-query remains stubbornly high for frontier models.

Industry sources quoted in the piece warn that the current pricing models are unsustainable and predict a wave of consolidation as smaller AI companies run out of runway. The analysis suggests that inference efficiency, not just model capability, will determine which AI companies survive the next phase of the market.