1 of 2
Story summary
- NVIDIA launched the Rubin CPX GPU for large-scale AI inference, processing over one million tokens for tasks like software coding and video generation.
- It offers 30 petaFLOPs of NVFP4 compute power and 128 GB of GDDR7 memory, enhancing performance during the context phase of inference.
- The Rubin CPX integrates into the Vera Rubin NVL144 CPX platform, combining 144 GPUs and 36 CPUs for 8 exaFLOPs of AI compute.
- Its architecture improves efficiency by separating context and generation phases, reducing latency.
- NVIDIA anticipates $5 billion returns for every $100 million invested in Rubin CPX infrastructure, with a release planned for late 2026.
1 / 2
