Microsoft is urging the AI industry to judge infrastructure by useful output rather than by the number of chips, tokens or datacenters it can deploy. The company calls this approach “yield,” borrowing the semiconductor industry’s term for the share of usable chips produced from a wafer.

The argument reflects a practical constraint: agentic systems consume much more compute than simple chat. Microsoft says a single agent task can use more than 3,400 times as many tokens as a typical chat interaction, while power, memory capacity and data movement are already limiting deployments. Long-running agents also need to retain more context close to the processors that use it.

Microsoft’s proposed response is to optimize the whole stack together. It points to model compression, better management of working memory, compilers that place data closer to compute, and networking designed around fleet-scale inference. This is a strategy statement rather than a new product launch, but it signals where Microsoft expects its AI infrastructure work to concentrate: extracting more practical work from each watt, byte and chip before simply adding more hardware.