Together AI launched Provisioned Throughput, giving customers reserved inference capacity for open models including MiniMax M3 and GLM-5.2. The service uses token-based pricing and includes a 99% uptime SLA.

The pitch is aimed at teams that want more predictable production inference without managing GPU-hour planning or infrastructure directly. Together says the model can also lower costs compared with proprietary API options.

For enterprises standardizing on open models, reserved throughput could make capacity planning and cost forecasting less volatile as agent and application traffic grows.