Microsoft has expanded Foundry Model Router from two regions to 28 for global-standard deployments and 21 for data-zone deployments. The wider footprint makes the routing service practical for more organizations that must keep inference within defined geographic boundaries.

The supported pool now includes Claude Opus 4.8 and the GPT-5.6 family, while four end-of-life models were removed. Default deployments receive pool updates automatically without a new endpoint or redeployment. Teams that restrict routing to a chosen subset must explicitly add new models, giving them more control over behavioral changes.

That distinction matters because a stable API does not guarantee stable output, latency, token use or tool selection. Microsoft offers balanced, quality and cost routing modes, but the effective context window is limited by the smallest model in the pool. Claude models must also be deployed separately in the same Foundry account with a matching service tier before the router can select them.

Responses identify the chosen model, allowing teams to audit routing after the fact. Microsoft has not published comparative accuracy, cost or latency figures for this expansion, and advises customers to benchmark before production use. Platform teams should therefore treat an automatic pool refresh like a dependency update, monitor the selected models and retain a restricted configuration as a fallback.