CoreWeave has made Nvidia Vera Rubin NVL72 systems available in its cloud, with coding-agent company Cognition as the first customer running production workloads. In early tests using tasks sampled from FrontierCode, Cognition reported up to 4.8 times the total token throughput of a GB200 NVL72 baseline for SWE-2 inference workloads.

The launch pairs the racks with Spectrum-X 102.4T Ethernet networking. Customers can operate the capacity through services including CoreWeave Kubernetes Service, Mission Control, Sandboxes and Inference. The benchmark is an early customer result on a particular software-engineering workload, not a general measure of performance across every model or application.

CoreWeave also plans to offer Nvidia’s Vera CPU for agent workloads. A rack contains 128 CPUs and 11,264 cores, and CoreWeave says its tests produced more than threefold faster sandbox startup and a 1.7-times gain across passing Terminal-Bench tasks.

Alongside the hardware, CoreWeave introduced Forge, a connected environment combining Weights & Biases, OpenPipe post-training tools and the marimo notebook project. The aim is to feed production traces back into evaluation and model improvement while keeping support open across models, frameworks and clouds.