Dell and Nvidia Introduced Hybrid AI Infrastructure
New workstation hardware and reference architecture aim to reduce high cloud-based token costs for enterprise AI agents.
Updated on Sept. 23, 2026 in Artificial Intelligence

Live Poll
Do you trust that businesses can keep the costs of adopting AI technologies under control?
Dell and Nvidia have introduced a hybrid AI workstation stack designed to manage agentic AI workloads locally rather than relying exclusively on public-cloud APIs. The system is designed to curb the unpredictable operating costs associated with rapidly rising enterprise token consumption.
Why it matters
Enterprises are seeking secure and economical ways to scale agentic AI, as the surge in reasoning usage has significantly increased cloud-based token expenses. This hybrid architecture allows firms to process routine tasks on local hardware while offloading complex frontier model requirements to the cloud.
The Dell Pro Max with GB10 workstation is powered by the Nvidia GB10 Grace Blackwell Superchip and features 128GB of unified memory. This hardware supports the NemoClaw reference stack, which uses OpenShell to isolate and govern agent activities.
The players
Dell
A multinational technology company providing servers, workstations, and integrated infrastructure stacks for enterprise data centers.
Nvidia
A semiconductor designer specializing in GPU-accelerated computing, AI superchips, and the software ecosystem for high-performance machine learning.
The details
The system employs a hybrid AI architecture that allows organizations to execute open models locally. To optimize efficiency, the NemoClaw reference stack incorporates OpenShell, a framework used to isolate and govern agent activity, ensuring that coordination agents can manage tasks using significantly fewer tokens than traditional large context window approaches.
Timeline
Two years ago, hyperscaler AI capital expenditure was estimated at USD 250 billion.
In 2026, hyperscaler AI capital expenditure surpassed USD 800 billion.
The Tech Race
This introduction follows a period of rapid industry consolidation where AI investment has shifted from experimental research to massive hyperscaler capital spending. Dell and Nvidia are attempting to shift this trajectory by moving the processing point closer to the enterprise end-user.
Organizations struggling with unpredictable cloud-based token costs can now evaluate hardware-based local processing for their internal agents. This shift requires adopting the NemoClaw reference stack to balance local model execution with necessary access to larger, cloud-hosted frontier models.
The takeaway
The move toward local agentic processing is a direct response to the massive growth in token-based cloud expenses. Watch for future performance benchmarks on the Dell Pro Max workstation to see if the reduction in token usage remains consistent when scaling to production-grade workloads.
Further reading
For more on the changing landscape of enterprise machine learning, explore our coverage of Artificial Intelligence.
Live Poll
Do you trust that businesses can keep the costs of adopting AI technologies under control?









