CoreWeave Launched Nvidia Vera Rubin NVL72 System
The GPU-accelerated rack system offers liquid-cooled infrastructure for large-scale AI model inference and development.
Updated on Sept. 30, 2026 in Data Centers

Live Poll
Do you trust companies to securely integrate AI systems into critical clinical workloads?
CoreWeave has made the Nvidia Vera Rubin NVL72 system available on its cloud platform, following successful initial deployments. The rack-scale architecture is designed to support intensive AI workloads through advanced cooling and modular design.
Why it matters
The platform introduces the Vera CPU, which is specifically engineered to remove infrastructure bottlenecks and accelerate agentic AI applications. This launch marks a significant expansion in the available compute density for enterprise-scale artificial intelligence.
The NVL72 rack-scale system integrates 36 CPUs and 72 GPUs, supported by NVLink 6 switches and BlueField-4 DPUs. It achieves higher efficiency through 100 percent liquid cooling and cable-free modular tray designs.
The players
CoreWeave
A specialized cloud infrastructure provider focused on high-performance compute and large-scale GPU clusters for AI workloads.
Nvidia
A semiconductor company that designs the GPU, CPU, and networking stack powering the Vera Rubin rack-scale platform.
Cognition
An AI research lab building autonomous software agents that served as early adopters for the Vera Rubin system.
The details
The NVL72 system organizes compute resources into high-density racks containing up to 128 Vera CPUs and 11,264 cores. Vera CPUs act as the primary engines for agentic AI—autonomous software agents capable of executing multi-step tasks—while ConnectX-9 SuperNICs manage high-speed networking between racks. CoreWeave also launched Forge, a software-defined development layer, to enable customers to build and evaluate models directly within the cloud environment.
Timeline
Early September 2026: Cognition began using the Vera Rubin systems.
September 30, 2026: CoreWeave announced the general availability of the Vera Rubin platform and Forge.
Coming weeks: Customers are expected to begin testing the standalone Vera CPU bare-metal offering.
The Tech Race
The Vera Rubin system follows the industry-standard rack-scale architecture established by the Nvidia GB200 NVL72. It represents a direct evolution of high-density data center design, aiming to maintain leadership in throughput for the most demanding inference workloads.
Enterprises and developers currently facing latency bottlenecks with AI model inference can expect significantly higher throughput from these systems. Access is provided through the CoreWeave cloud platform, with standalone bare-metal testing expected to begin shortly.
The takeaway
The Vera Rubin architecture demonstrates a pivot toward specialized, agent-focused CPU hardware in high-density rack form factors. Watch for the forthcoming bare-metal testing phase to gauge how these systems perform outside of the managed cloud-platform environment.
What happens next
Customers are scheduled to begin testing the standalone Vera CPU bare-metal offering in the coming weeks.
Further reading
For more on the scaling of large-scale infrastructure, visit Data Centers.
Live Poll
Do you trust companies to securely integrate AI systems into critical clinical workloads?










