What Is the NVIDIA H200 NVL?
The NVIDIA H200 NVL is the PCIe-based form factor of the H200 GPU, built on the Hopper architecture. It delivers the same core memory capability as the SXM5 version 141GB of HBM3e memory at 4.8TB/s but in a dual-slot, air-cooled PCIe Gen5 card that fits into standard enterprise servers.
For UK organisations that want to add GPU capability without committing to a full HGX baseboard, the H200 NVL is the more flexible option. It slots into HPE ProLiant, Dell PowerEdge, Supermicro, and Lenovo ThinkSystem platforms, giving infrastructure teams the freedom to build GPU capacity at their own pace.
The NVL variant supports 2-way and 4-way NVLink bridge configurations, enabling GPU-to-GPU communication at up to 900GB/s per GPU. This makes it suitable for large language model inference, HPC simulation, and AI inference workloads that benefit from pooled GPU memory.
H200 NVL vs H200 HGX: What Is the Difference?
The H200 ships in two form factors: the NVL (PCIe) and the HGX (SXM5). They share the same GPU die and memory, but the packaging, power envelope, and server integration differ significantly.
| Specification | H200 NVL (PCIe) | H200 HGX (SXM5) |
|---|---|---|
| Form factor | PCIe Gen5 dual-slot, air-cooled | SXM5 module on HGX baseboard |
| GPU memory | 141GB HBM3e | 141GB HBM3e |
| Memory bandwidth | 4.8 TB/s | 4.8 TB/s |
| TDP | Up to 600W (configurable) | Up to 700W (configurable) |
| Interconnect | NVLink bridge (2 or 4-way, 900GB/s per GPU) | HGX NVLink backplane (900GB/s per GPU) |
| PCIe bandwidth | 128 GB/s (Gen5) | 128 GB/s (Gen5) |
| MIG instances | Up to 7 @ 16.5GB each | Up to 7 @ 18GB each |
| Cooling | Air-cooled (standard) | Air or liquid (typically liquid at scale) |
| Server integration | Standard PCIe slots in enterprise servers | HGX baseboard (4 or 8 GPUs) |
| NVIDIA AI Enterprise | Included (5-year) | Add-on |
| Typical deployment | 1-8 GPUs per server, incremental scaling | 4-8 GPUs per server, dense cluster |
The practical difference comes down to deployment. The H200 NVL fits into existing server infrastructure and scales incrementally. The H200 HGX is denser, typically liquid-cooled, and designed for clusters of 4 or 8 GPU baseboards. For UK enterprises that want to start with one or two GPUs and grow, the NVL is the natural choice.
Technical Specifications
| Specification | H200 NVL |
|---|---|
| Architecture | NVIDIA Hopper (GH100) |
| GPU memory | 141GB HBM3e |
| Memory bandwidth | 4.8 TB/s |
| FP8 Tensor Core | 3,958 TFLOPS (with sparsity) |
| FP16 / BF16 Tensor Core | 1,979 TFLOPS (with sparsity) |
| TF32 Tensor Core | 989 TFLOPS (with sparsity) |
| FP32 | 60 TFLOPS |
| FP64 Tensor Core | 60 TFLOPS |
| FP64 | 30 TFLOPS |
| INT8 Tensor Core | 3,958 TFLOPS (with sparsity) |
| TDP | Up to 600W (configurable) |
| Form factor | PCIe Gen5, dual-slot, air-cooled |
| NVLink interconnect | 900 GB/s per GPU (via bridge) |
| PCIe bandwidth | 128 GB/s |
| NVLink bridge | 2-way or 4-way |
| MIG instances | Up to 7 @ 16.5 GB each |
| Video decoders | 7 NVDEC, 7 JPEG |
| Confidential computing | Supported |
| NVIDIA AI Enterprise | Included (5-year subscription) |
The H200 NVL and H200 HGX share identical compute performance and memory. The FP8 Tensor Core throughput of 3,958 TFLOPS is the same across both form factors. The only hardware differences are TDP, MIG partition sizes, and the cooling and interconnect implementation.
NVLink Bridge Configurations for H200 NVL
The H200 NVL supports two NVLink bridge configurations. This is the key differentiator from the SXM form factor: instead of a fixed backplane interconnect, the NVL uses physical bridge cards that connect pairs or quads of GPUs.
| Configuration | GPUs | NVLink bandwidth | Use case |
|---|---|---|---|
| 2-way NVLink | 2 x H200 NVL | 900 GB/s per GPU | Single-node LLM inference, dual-GPU training |
| 4-way NVLink | 4 x H200 NVL | 900 GB/s per GPU | Multi-GPU inference, larger model serving, HPC |
| No bridge | 1 x H200 NVL | N/A | Single-GPU inference, dev/test, edge deployment |
In a 4-way configuration, the H200 NVL GPUs share a unified memory pool of 564GB, which is sufficient to run large language models like Llama 2 70B entirely in GPU memory.
The NVIDIA Enterprise Reference Architecture for H200 NVL recommends a minimum of 4 GPUs per node for production AI workloads, with Spectrum-X Ethernet networking and NCCL for multi-node scaling.
Server Platforms and Compatibility
The H200 NVL fits into standard PCIe Gen5 server platforms from all major OEMs. Because it uses the PCIe form factor, compatibility is broader than the SXM5 module.
| OEM | Example Platform | Max H200 NVL GPUs |
|---|---|---|
| HPE | ProLiant DL380a Gen11 / DL385 Gen11 | Up to 4 (with NVLink bridge) |
| Dell | PowerEdge XE9680 / R760xa | Up to 4 (with NVLink bridge) |
| Supermicro | AS-4125GS-TNRT2 / SYS-421GU-TNRT2 | Up to 8 |
| Lenovo | ThinkSystem SR680a V3 | Up to 4 (with NVLink bridge) |
| NVIDIA MGX | Partner reference architecture | Up to 8 |
Eton Technology supplies compatible HPE and Dell server platforms that support H200 NVL deployment, including the DL380a Gen11 and PowerEdge XE9680.
Browse Compatible Server Platforms
Deploying H200 NVL in UK Data Centres
Deploying H200 NVL GPUs in a UK data centre environment requires planning around power, cooling, and network infrastructure.
Power. Each H200 NVL has a configurable TDP up to 600W. A 4-GPU node draws up to 2.4kW for the GPUs alone, plus the host server. Standard UK data centre racks at 8-12kW per rack can hold 3-4 such nodes.
Cooling. The NVL is air-cooled, which simplifies deployment in existing facilities. For dense deployments (8 GPUs per node), supplemental fan kits or rear-door heat exchangers may be needed.
Networking. For multi-node scaling, NVIDIA recommends Spectrum-X Ethernet networking at 400GbE or higher. NVLink handles intra-node GPU communication, and NCCL manages inter-node communication over standard Ethernet or InfiniBand.
Lead time. H200 NVL availability has improved significantly since launch. Contact the Eton Technology team for current pricing and lead times.
H200 NVL for AI Inference and Training
The H200 NVL is designed primarily for AI inference, though it handles training workloads well within its memory and bandwidth envelope.
| Workload | Suitable config | Notes |
|---|---|---|
| LLM inference (Llama 2 70B, GPT-3 class) | 2-4 way NVLink | Fits entirely in 2-4 GPU memory pool |
| LLM fine-tuning | 4-way NVLink | LoRA/QLoRA on single GPU, full fine-tune on 4 |
| HPC simulation (MILC, NAMD, GROMACS) | 2-4 way NVLink | Benefits from 4.8TB/s memory bandwidth |
| RAG and embedding pipelines | Single GPU | Embedding models fit in 141GB easily |
| Batch inference (multi-tenant) | Single GPU with MIG | Partition into 7 instances for isolation |
| Video transcoding and processing | Single GPU | 7 NVDEC decoders handle heavy throughput |
NVIDIA reports up to 1.7x faster LLM inference on the H200 NVL compared to the H100 NVL, driven by the larger 141GB memory pool and higher 4.8TB/s bandwidth. For UK enterprises running production inference, this translates to lower latency, higher throughput, and fewer GPUs needed to serve the same model.
Frequently Asked Questions
What does NVL stand for in NVIDIA GPUs?
NVL stands for NVLink. It indicates the GPU variant supports NVLink bridge connectivity, allowing multiple GPUs to share a high-speed interconnect for pooled memory and accelerated multi-GPU workloads.
What is the difference between H200 NVL and H200 HGX?
The H200 NVL is a PCIe dual-slot air-cooled card that fits into standard servers. The H200 HGX uses the SXM5 form factor and sits on a dedicated HGX baseboard. The GPU memory and compute performance are identical, but the NVL has a lower TDP (600W vs 700W) and includes a 5-year NVIDIA AI Enterprise subscription.
Can I use H200 NVL in an existing server?
Yes, provided the server has PCIe Gen5 slots with sufficient power delivery (600W per GPU) and physical clearance for dual-slot cards. HPE ProLiant, Dell PowerEdge, Supermicro, and Lenovo ThinkSystem platforms all support H200 NVL.
How many H200 NVL GPUs can I connect with NVLink?
The H200 NVL supports 2-way and 4-way NVLink bridge configurations. In a 4-way setup, four GPUs share a unified 564GB memory pool with 900GB/s bi-directional bandwidth per GPU.
Does the H200 NVL include NVIDIA AI Enterprise?
Yes. The H200 NVL comes with a 5-year NVIDIA AI Enterprise subscription included in the price. The H200 HGX (SXM5) requires AI Enterprise as a separate add-on purchase.
What is the lead time for H200 NVL in the UK?
Availability varies. Eton Technology can provide current lead time estimates and available configurations. Contact the team for up-to-date pricing and stock information.
Does Eton Technology supply H200 NVL GPUs to UK enterprises?
Yes. Eton Technology supplies H200 NVL GPUs, H200 HGX baseboards, and compatible HPE, Dell, and Lenovo server platforms to UK enterprises, research institutions, and government organisations.
Get in Touch
Eton Technology Solutions is a UK-based supplier of NVIDIA enterprise GPU infrastructure. The team supplies H200 NVL GPUs, H200 HGX baseboards, and compatible server platforms from HPE, Dell, Lenovo, and Supermicro.
Browse NVIDIA GPUs Shop H200 NVL
Phone: (+44) 333 999 7768 | Email: info@etontechnology.com
