Full Breakdown
MLPerf Inference v5.1: A Benchmark Showdown Among AI Hardware Giants
9/10/2025, 1:10:38 PM
Overview of MLPerf Inference v5.1 Results
The latest MLPerf Inference v5.1 benchmarks, released by MLCommons, highlight significant advancements in AI inference performance across various hardware platforms, particularly focusing on NVIDIA, AMD, and Intel. This round of benchmarks included submissions from 27 organizations, showcasing new models and technologies designed to meet the increasing demands of AI workloads.
Key Performance Highlights
NVIDIA's Blackwell Ultra GB300 system achieved record-breaking performance, setting new benchmarks in multiple categories, including the DeepSeek-R1 reasoning model and the Llama 3.1 series. The GB300 demonstrated up to a 45% increase in throughput compared to its predecessor, the GB200, and maintained a significant lead over competing architectures. For instance, it recorded 5,842 tokens per second in the DeepSeek-R1 offline benchmark, outperforming the previous generation by a substantial margin.
In contrast, AMD's Instinct MI355X also made notable strides, achieving a 2.7x performance improvement in token generation compared to earlier models. The MI355X demonstrated competitive results, particularly in the Llama 2 70B benchmark, where it generated up to 648,248 tokens per second in a 64-chip configuration.
Intel's Arc Pro B60, part of the Project Battlematrix initiative, showcased its capabilities with a performance per dollar advantage of up to 1.25x over NVIDIA's RTX Pro 6000. The Arc Pro B60 is designed to cater to high-end workstations and edge applications, emphasizing accessibility and ease of use.
Technological Innovations Driving Performance
NVIDIA's success in the MLPerf v5.1 benchmarks can be attributed to several technological innovations, including the use of the NVFP4 data format for improved accuracy and efficiency. The Blackwell Ultra architecture incorporates advanced parallelism techniques, allowing for disaggregated serving, which optimizes the context and generation phases of AI inference across separate GPUs. This approach significantly enhances throughput and reduces latency.
Intel's new inference-optimized software stack for the Arc Pro B-Series GPUs aims to simplify the deployment of AI applications while ensuring enterprise-class reliability. The integration of multi-GPU scaling and PCIe P2P data transfers further enhances the performance of Intel's offerings.
Criticism and Competitive Landscape
Despite NVIDIA's dominance, the competition is intensifying. AMD's increased participation in MLPerf benchmarks and the introduction of the MI355X indicate a strategic push to capture market share in the AI hardware space. Critics argue that while NVIDIA leads in performance, the high resource demands of running MLPerf benchmarks may deter other companies from participating, potentially skewing the competitive landscape.
Official Statements
Lisa Pearce, Intel's corporate vice president, stated, “The MLPerf v5.1 results are a powerful validation of Intel’s GPU and AI strategy,” highlighting the significance of their new offerings in the evolving AI landscape. Meanwhile, NVIDIA emphasized the importance of their Blackwell Ultra architecture in achieving record performance, showcasing their commitment to innovation in AI hardware.
What's Next for AI Inference Hardware?
As the demand for AI inference capabilities continues to grow, future MLPerf submissions are expected to reflect ongoing optimizations from NVIDIA, AMD, and Intel. The competitive dynamics in this space will likely drive further advancements in hardware and software, enhancing the overall efficiency and effectiveness of AI applications across various industries.
