MLPerf Inference v6.1 brought a record 30 submitters and 486 datacentre and edge results, the first peer reviewed Vera Rubin NVL72 numbers, and a 512 GPU AMD run. The useful signal for buyers is how much of the gain is software rather than new silicon.
What was announced
- MLCommons published MLPerf Inference v6.1 with 30 submitting organisations and 486 datacenter and edge results, a participation record.
- The round adds two tests: an End-to-End RAG pipeline for the datacenter and an Edge Agentic Inference benchmark for single user devices.
- MLCommons says the best per accelerator DeepSeek-R1 result in the server scenario is 5.7x better than in v5.1 a year ago, and the best VLM result improved 2.99x in the six months since v6.0.
- NVIDIA's Vera Rubin NVL72 appears in the preview category, submitted by NVIDIA and by Nebius on its VR200 NVL72, meaning it is expected to be commercially available by the next round.
- NVIDIA reports up to 2.5x higher token throughput than GB300 NVL72 on DeepSeek-R1 using TensorRT-LLM, and up to 3.7x on the Qwen3 vision language model using vLLM with NVIDIA Dynamo, both against its own prior generation.
- AMD submitted six model families using Instinct MI355X, MI350X and the new MI350P PCIe card.
- Crusoe ran the largest system in MLPerf Inference history with AMD: 512 Instinct MI355X GPUs across 64 nodes on a standard RoCE Ethernet fabric, reporting 5.75 million tokens per second offline and 5.39 million in server on gpt-oss-120b, and 2.90 million offline and 2.41 million server on DeepSeek-R1, with throughput scaling near linearly from 1 to 64 nodes.
- Intel ran a four GPU Arc Pro B70 node with 128GB combined VRAM and reports gpt-oss-120B improved 36% in server and 27% in offline over v6.0 on the same hardware.
- Intel says Xeon 6980P Llama 3.1-8B server throughput rose 2.4x from v6.0 on identical silicon, a software only gain.
The ETON view
Two numbers in this round matter more than the headline leadership claims. Intel's 2.4x Xeon 6980P gain and its 36% Arc Pro B70 gain both came on unchanged hardware, and NVIDIA's own year on year per accelerator figure includes a large software component. If your inference platform is 12 to 24 months old, a stack upgrade is very likely the cheapest capacity you can buy this quarter, and it should be tested before any refresh business case is signed.
The Crusoe result is the second useful signal: 512 MI355X GPUs scaling near linearly across 64 nodes on standard RoCE Ethernet, not a proprietary fabric. For hosting providers and AI platforms sourcing in the UK, that means the fabric decision no longer forces a single vendor path, and Ethernet based clusters can be built from parts that are actually available at a known lead time.
Vera Rubin NVL72 is still preview, so it is a roadmap item rather than a procurement option, and NVL72 class density needs liquid cooling and rack power that most UK colocation halls at 6 to 12 kW per rack cannot deliver without a contract change. The practical near term plan for most buyers is Hopper and Blackwell class capacity at current market pricing, sized on measured tokens per second for their own models, with the newest platforms treated as a planned second phase. We source new and certified refurbished GPU platforms across vendors, quote on lead time as well as price, and will say when the honest answer is to keep and retune what you already own.
Related infrastructure
Category: AI & GPU · Vendor: NVIDIA, AMD, Intel · Technology: Vera Rubin NVL72, Instinct MI355X, MI350P, Arc Pro B70, Xeon 6980P · Last verified: 18 Sep 2026 · ~3 min read
Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.
