Drooid Logo
Back to story perspectives

Full Breakdown

NVIDIA Shows DSX AI Factory Platform Boosts Token Throughput and Enables Grid-Responsive Operations

By Drooid · · How we work

Core Event: Early Production Results Demonstrate 24% Token-Throughput Gain

On September 15, 2026, NVIDIA released early production results for its DSX AI factory platform. In a validation run by cloud-provider Lambda, a five-rack, 19-node cluster running DSX MaxLPS delivered 24 % more cluster-wide token throughput—rising from roughly 4 million to 5 million tokens per second—while staying within the same power budget as a 16-node configuration. NVIDIA reported a 23 % improvement in performance-per-watt for the same workload. The results were highlighted in a keynote by Ian Buck, NVIDIA’s vice-president of hyperscale and high-performance computing, at the AI Infra Summit in Santa Clara.

Background & Context: DSX Platform Evolution and Power Architecture

NVIDIA announced the DSX platform at GTC Taipei on May 31 2026. The suite combines open-source libraries, APIs, reference designs, and NVIDIA hardware into a unified framework for AI-factory design, deployment, and operation. At launch, cloud partners CoreWeave, Crusoe, Firmus, IREN, Lambda, Nebius, Nscale, and Yotta Data Services were deploying DSX Sim, MaxLPS, and OS, while OEMs Dell, HPE, Lenovo, and Supermicro were building DSX-ready systems.

The platform uses an 800 VDC power architecture to reduce conversion losses. NVIDIA notes that a GB200 NVL72 rack with direct-liquid cooling dissipates roughly 120 kW of heat before reaching compute units, highlighting the need for precise power management.

Data & Statistics: Quantitative Highlights

Data & Statistics: Quantitative Highlights
MetricFigureSource
Token throughput (baseline)~4 million tokens / secNVIDIA
Token throughput (DSX MaxLPS)~5 million tokens / secNVIDIA
Performance-per-watt gain23 %NVIDIA
Power reduction during demand response (August)4 MW -> 3 MWNVIDIA
Demand-signal events received>200 (as of Sep 2026)NVIDIA
Data-center revenue Q2 FY2027$89.02 billion (117 % YoY)NVIDIA
Vera Rubin chips share Q3 FY2027 revenue~20 %NVIDIA
Cost per gigawatt (Vera Rubin)$40 billion per GWJensen Huang (paraphrased)

Official Statements & Responses

Varun Sivaram, founder and CEO of Emerald AI, described the August demand-response test as the company’s first deployment across thousands of NVIDIA GPUs, noting that the Conductor platform automatically shed low-priority jobs while preserving high-priority inference workloads.

Nico Procos, senior vice-president at Silicon Valley Power, said the pilot would evaluate tools to protect grid reliability and affordability while supporting flexible planning for future load growth.

On-the-Ground Reports: Demand-Response in Action

During an August heat-wave, Silicon Valley Power sent a grid-signal to NVIDIA’s Eos AI factory in Santa Clara. Emerald AI’s Conductor platform executed a predefined hierarchy: low-priority jobs yielded, high-priority inference continued, and overall power consumption dropped from four megawatts to three megawatts in under a minute, without human intervention. Since the pilot’s announcement on April 21 2026, the utility has issued more than 200 demand signals, each met within a minute.

What’s Next

The first dedicated DSX Flex commercial deployment—a 96-megawatt Vera Rubin AI factory at NVIDIA’s AI Factory Research Center in Manassas, Virginia—is slated for later in 2026, in collaboration with Digital Realty, EPRI, and PJM Interconnection. The AI Infra Summit will continue from September 15–17 2026 at the Santa Clara Convention Center, where further demonstrations of DSX capabilities are expected.