Full Breakdown
NVIDIA Shows DSX AI Factory Platform Boosts Token Throughput and Enables Grid-Responsive Operations
By Drooid · · How we work
Core Event: Early Production Results Demonstrate 24% Token-Throughput Gain
On September 15, 2026, NVIDIA released early production results for its DSX AI factory platform. In a validation run by cloud-provider Lambda, a five-rack, 19-node cluster running DSX MaxLPS delivered 24 % more cluster-wide token throughput—rising from roughly 4 million to 5 million tokens per second—while staying within the same power budget as a 16-node configuration. NVIDIA reported a 23 % improvement in performance-per-watt for the same workload. The results were highlighted in a keynote by Ian Buck, NVIDIA’s vice-president of hyperscale and high-performance computing, at the AI Infra Summit in Santa Clara.
Background & Context: DSX Platform Evolution and Power Architecture
NVIDIA announced the DSX platform at GTC Taipei on May 31 2026. The suite combines open-source libraries, APIs, reference designs, and NVIDIA hardware into a unified framework for AI-factory design, deployment, and operation. At launch, cloud partners CoreWeave, Crusoe, Firmus, IREN, Lambda, Nebius, Nscale, and Yotta Data Services were deploying DSX Sim, MaxLPS, and OS, while OEMs Dell, HPE, Lenovo, and Supermicro were building DSX-ready systems.
The platform uses an 800 VDC power architecture to reduce conversion losses. NVIDIA notes that a GB200 NVL72 rack with direct-liquid cooling dissipates roughly 120 kW of heat before reaching compute units, highlighting the need for precise power management.
Data & Statistics: Quantitative Highlights
| Metric | Figure | Source |
|---|---|---|
| Token throughput (baseline) | ~4 million tokens / sec | NVIDIA |
| Token throughput (DSX MaxLPS) | ~5 million tokens / sec | NVIDIA |
| Performance-per-watt gain | 23 % | NVIDIA |
| Power reduction during demand response (August) | 4 MW -> 3 MW | NVIDIA |
| Demand-signal events received | >200 (as of Sep 2026) | NVIDIA |
| Data-center revenue Q2 FY2027 | $89.02 billion (117 % YoY) | NVIDIA |
| Vera Rubin chips share Q3 FY2027 revenue | ~20 % | NVIDIA |
| Cost per gigawatt (Vera Rubin) | $40 billion per GW | Jensen Huang (paraphrased) |
Official Statements & Responses
Varun Sivaram, founder and CEO of Emerald AI, described the August demand-response test as the company’s first deployment across thousands of NVIDIA GPUs, noting that the Conductor platform automatically shed low-priority jobs while preserving high-priority inference workloads.
Nico Procos, senior vice-president at Silicon Valley Power, said the pilot would evaluate tools to protect grid reliability and affordability while supporting flexible planning for future load growth.
On-the-Ground Reports: Demand-Response in Action
During an August heat-wave, Silicon Valley Power sent a grid-signal to NVIDIA’s Eos AI factory in Santa Clara. Emerald AI’s Conductor platform executed a predefined hierarchy: low-priority jobs yielded, high-priority inference continued, and overall power consumption dropped from four megawatts to three megawatts in under a minute, without human intervention. Since the pilot’s announcement on April 21 2026, the utility has issued more than 200 demand signals, each met within a minute.
What’s Next
The first dedicated DSX Flex commercial deployment—a 96-megawatt Vera Rubin AI factory at NVIDIA’s AI Factory Research Center in Manassas, Virginia—is slated for later in 2026, in collaboration with Digital Realty, EPRI, and PJM Interconnection. The AI Infra Summit will continue from September 15–17 2026 at the Santa Clara Convention Center, where further demonstrations of DSX capabilities are expected.
