Cheaper AI output has so far increased demand faster than it has reduced infrastructure pressure. Data from Ornn, Silicon Data and Bloomberg through August 2026 shows token prices continuing to fall while rental prices for Nvidia H100 GPUs remain steady or increase.

The pattern resembles the Jevons paradox: when a resource becomes cheaper to use, new applications can expand total consumption. Lower token prices make long-running agents and automation more affordable, while those systems may generate many more model calls than an ordinary chat session.

That distinction complicates demand estimates. It is not clear how much growth comes from more people using AI and how much comes from agents repeatedly calling models, tools and other agents. A modest increase in users can still create a much larger increase in compute if each workflow consumes many steps.

The trend is favorable for chipmakers, cloud providers, memory suppliers and energy companies only while usage expands faster than unit costs fall. If demand growth slows, declining token prices could finally reach the hardware market and weaken utilization and rental rates. The current data documents a period of persistent compute demand; it does not guarantee that the relationship will continue.