RunPod’s Deploy When Available lets customers request a GPU spec even when capacity is unavailable and have the workload launch once resources open up. The company says the feature is now generally available.

For AI developers, this tackles a real infrastructure pain point: popular GPU types can be fully occupied, forcing teams to poll dashboards or build workarounds. Queueing turns that scarcity into a managed workflow.

The update fits a broader trend in AI infrastructure, where scheduling and availability tooling are becoming as important as raw GPU access.