Modal has made its multi-node GPU cluster service generally available to all workspaces. Developers can turn a function into a clustered workload with one decorator, request several GPU nodes together and enable remote direct memory access, or RDMA, with a flag.
The platform “gang schedules” every node required by a job at once, avoiding partial clusters that cannot begin work. Modal says clusters draw from the same shared capacity pool as its other workloads, can start within seconds and are billed by the second rather than through an hourly reservation.
RDMA lets network adapters move data directly between GPU memory on different hosts without routing it through host memory. Modal reports up to 6.4 terabits per second per node and automatically configures PyTorch and the NCCL communications library. The company also added RDMA support to its preferred gVisor container runtime and contributed those changes upstream.
Clusters integrate with Modal volumes, cloud-bucket mounts and queues. Customers are already using them for large-model fine-tuning, robotics research and distributed video-model inference. Access is broad, but the maximum cluster size still depends on each workspace’s GPU limits, so unusually large jobs require coordination with Modal.