CoreWeave Expanded Stack Beyond GPU Compute
The company launched a new software layer and networking suite to meet shifting AI infrastructure demands.
Updated on Sept. 30, 2026 in Data Centers

Live Poll
Do you trust that new AI infrastructure investments will ultimately lower costs for everyday users?
CoreWeave has announced an expansion of its infrastructure stack, moving beyond GPU compute to include integrated networking, storage, and a new software layer called Forge. This development aims to address the growing market shift toward inference-heavy AI workloads.
Why it matters
AI infrastructure optimization is evolving from a focus on individual components to system-level performance, driven by a need to lower the cost of producing tokens at scale. This shift reflects a market transition where inference demand is now outpacing model training requirements.
The Forge system integrates a closed-loop runtime environment that provides observability and security features alongside networking and storage. It is designed to optimize system-level performance as the ratio of training to inference workloads is projected to shift from 50/50 to 10/90 by 2027.
The players
CoreWeave
A cloud provider specializing in GPU-accelerated infrastructure for AI training and large-scale model inference.
The details
CoreWeave expanded its platform to support full-stack AI deployment, moving past raw GPU access to provide orchestration across networking and storage hardware. The newly introduced Forge system acts as a central control plane for runtime curation, allowing users to monitor performance and security in real time. By integrating evaluation and observation directly into the infrastructure, the stack aims to help users reduce the cost per token for inference, which is increasingly prioritized over training speed.
Timeline
September 30, 2026: CoreWeave announced the Forge software layer at the Fully Connected event.
2027: The industry is expected to reach a 10/90 ratio for training-to-inference workloads.
The Tech Race
The expansion signals a departure from pure GPU-rental models toward vertically integrated systems designed for production-grade AI. This move tracks the industry-wide shift toward optimized inference, following the documented pattern of companies reconfiguring their hardware stacks to handle production workloads.
Customers seeking more flexible AI infrastructure can now access shorter contract terms and on-demand pricing models via the new stack. These changes primarily benefit enterprises and developers scaling production inference workloads who require higher system-level efficiency.
The takeaway
The rapid rise of inference-heavy workloads is forcing providers to integrate security and observability directly into their underlying compute layers. Keep an eye on the 10/90 training-to-inference workload ratio in 2027 to see if cost-per-token improvements manifest as predicted.
Further reading
For more on how infrastructure providers are adjusting to capacity needs, see our Data Centers section.
Live Poll
Do you trust that new AI infrastructure investments will ultimately lower costs for everyday users?










