RunPod has detailed a private-pool model that reserves GPU capacity for one organization while allowing eligible workloads to burst into shared on-demand supply. The same reserved pool can support Pods, serverless endpoints and clusters, depending on the customer agreement.
Unlike usage-based instances, reserved GPUs are billed throughout the commitment whether they are active or idle. RunPod therefore recommends sizing a pool around the workload’s predictable floor rather than its highest spike. A service that uses four GPUs most of the day and occasionally reaches 20 can reserve the steady four and send temporary demand to on-demand machines when bursting is enabled. Extra capacity is billed separately and is still subject to availability.
The arrangement provides dedicated capacity inside RunPod’s managed container environment, not bare-metal access. Stopping a workload releases the job but not the underlying reservation. Teams can share that held capacity across daytime inference, overnight evaluations and fine-tuning to improve utilization. Hard caps and stop behavior for overflow are still being finalized and may depend on API support, so customers need to confirm whether each workload should burst, queue or fail. Experimental and unpredictable projects remain better suited to on-demand GPUs.