NVIDIA has extended NVLink Fusion with NVHBM, moving its own memory controller into the HBM base die. Amazon's Annapurna Labs is first in line. Here is why the memory layer, not the GPU, is now the interesting part of an AI platform decision.
What was announced
- NVIDIA has expanded NVLink Fusion with NVHBM, a high-bandwidth memory technology that integrates NVIDIA's custom memory controller into the HBM base die rather than the XPU die.
- NVIDIA states NVHBM delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption, and frees up to 25% more area on the XPU compute die, compared with standard HBM4E.
- NVIDIA is establishing a standard NVHBM implementation available from multiple memory providers, to reduce the effort of qualifying memory across suppliers.
- Amazon's Annapurna Labs will be the first to work on NVHBM, and will support NVLink Fusion starting with its Trainium4 generation.
- NVLink Fusion lets partners connect custom XPUs and CPUs to NVIDIA's rack-scale platform using NVLink chiplets, NVLink-C2C, NVLink Switches and MGX systems and racks.
- StorageReview reports NVIDIA claims up to a 67% reduction in PHY and peripheral support area versus standard HBM4E implementations.
The ETON view
Read past the silicon detail and this is a statement about where value sits in an AI platform. When the memory subsystem is worth 30% bandwidth and 15% power on its own, memory stops being a line item on a quote and becomes an architectural decision, and NVIDIA has just made sure it is present in that decision even on chips it does not make.
For the buyers we work with, two things follow. First, if you are specifying an AI platform for delivery in 2027 or later, ask the vendor what memory generation and configuration is actually committed, not just the accelerator name. Memory supply is the constraint in this market and configurations are being changed late. Second, the practical near-term effect of hyperscalers absorbing leading-edge memory is that everything below the frontier gets tighter and more expensive to buy at short notice. If you have a GPU or high-memory server requirement landing in the next two quarters, lock the specification and the supply now rather than at PO stage. That is a sourcing problem before it is a technology problem, and it is the part we can actually shorten.
Related infrastructure
Category: AI & GPU · Vendor: NVIDIA, AWS · Technology: NVHBM, NVLink Fusion, HBM4E, MGX · Last verified: 02 Sep 2026 · ~2 min read
Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.
