Skip to content
Eton Technology SolutionsEton Technology Solutions
NVIDIA H200 NVL GPU: Technical Guide for UK AI Infrastructure Teams

NVIDIA H200 NVL GPU: Technical Guide for UK AI Infrastructure Teams

July 28, 2026

NVIDIA H200 NVL GPU: Technical Guide for UK AI Infrastructure Teams

What Is the NVIDIA H200 NVL?

The NVIDIA H200 NVL is the PCIe-based form factor of the H200 GPU, built on the Hopper architecture. It delivers the same core memory capability as the SXM5 version 141GB of HBM3e memory at 4.8TB/s but in a dual-slot, air-cooled PCIe Gen5 card that fits into standard enterprise servers.

For UK organisations that want to add GPU capability without committing to a full HGX baseboard, the H200 NVL is the more flexible option. It slots into HPE ProLiant, Dell PowerEdge, Supermicro, and Lenovo ThinkSystem platforms, giving infrastructure teams the freedom to build GPU capacity at their own pace.

The NVL variant supports 2-way and 4-way NVLink bridge configurations, enabling GPU-to-GPU communication at up to 900GB/s per GPU. This makes it suitable for large language model inference, HPC simulation, and AI inference workloads that benefit from pooled GPU memory.

Shop H200 NVL

H200 NVL vs H200 HGX: What Is the Difference?

The H200 ships in two form factors: the NVL (PCIe) and the HGX (SXM5). They share the same GPU die and memory, but the packaging, power envelope, and server integration differ significantly.

Specification H200 NVL (PCIe) H200 HGX (SXM5)
Form factor PCIe Gen5 dual-slot, air-cooled SXM5 module on HGX baseboard
GPU memory 141GB HBM3e 141GB HBM3e
Memory bandwidth 4.8 TB/s 4.8 TB/s
TDP Up to 600W (configurable) Up to 700W (configurable)
Interconnect NVLink bridge (2 or 4-way, 900GB/s per GPU) HGX NVLink backplane (900GB/s per GPU)
PCIe bandwidth 128 GB/s (Gen5) 128 GB/s (Gen5)
MIG instances Up to 7 @ 16.5GB each Up to 7 @ 18GB each
Cooling Air-cooled (standard) Air or liquid (typically liquid at scale)
Server integration Standard PCIe slots in enterprise servers HGX baseboard (4 or 8 GPUs)
NVIDIA AI Enterprise Included (5-year) Add-on
Typical deployment 1-8 GPUs per server, incremental scaling 4-8 GPUs per server, dense cluster

The practical difference comes down to deployment. The H200 NVL fits into existing server infrastructure and scales incrementally. The H200 HGX is denser, typically liquid-cooled, and designed for clusters of 4 or 8 GPU baseboards. For UK enterprises that want to start with one or two GPUs and grow, the NVL is the natural choice.

Shop H200 NVL Shop HGX H200

Technical Specifications

Specification H200 NVL
Architecture NVIDIA Hopper (GH100)
GPU memory 141GB HBM3e
Memory bandwidth 4.8 TB/s
FP8 Tensor Core 3,958 TFLOPS (with sparsity)
FP16 / BF16 Tensor Core 1,979 TFLOPS (with sparsity)
TF32 Tensor Core 989 TFLOPS (with sparsity)
FP32 60 TFLOPS
FP64 Tensor Core 60 TFLOPS
FP64 30 TFLOPS
INT8 Tensor Core 3,958 TFLOPS (with sparsity)
TDP Up to 600W (configurable)
Form factor PCIe Gen5, dual-slot, air-cooled
NVLink interconnect 900 GB/s per GPU (via bridge)
PCIe bandwidth 128 GB/s
NVLink bridge 2-way or 4-way
MIG instances Up to 7 @ 16.5 GB each
Video decoders 7 NVDEC, 7 JPEG
Confidential computing Supported
NVIDIA AI Enterprise Included (5-year subscription)

The H200 NVL and H200 HGX share identical compute performance and memory. The FP8 Tensor Core throughput of 3,958 TFLOPS is the same across both form factors. The only hardware differences are TDP, MIG partition sizes, and the cooling and interconnect implementation.

NVLink Bridge Configurations for H200 NVL

The H200 NVL supports two NVLink bridge configurations. This is the key differentiator from the SXM form factor: instead of a fixed backplane interconnect, the NVL uses physical bridge cards that connect pairs or quads of GPUs.

Configuration GPUs NVLink bandwidth Use case
2-way NVLink 2 x H200 NVL 900 GB/s per GPU Single-node LLM inference, dual-GPU training
4-way NVLink 4 x H200 NVL 900 GB/s per GPU Multi-GPU inference, larger model serving, HPC
No bridge 1 x H200 NVL N/A Single-GPU inference, dev/test, edge deployment

In a 4-way configuration, the H200 NVL GPUs share a unified memory pool of 564GB, which is sufficient to run large language models like Llama 2 70B entirely in GPU memory.

The NVIDIA Enterprise Reference Architecture for H200 NVL recommends a minimum of 4 GPUs per node for production AI workloads, with Spectrum-X Ethernet networking and NCCL for multi-node scaling.

Server Platforms and Compatibility

The H200 NVL fits into standard PCIe Gen5 server platforms from all major OEMs. Because it uses the PCIe form factor, compatibility is broader than the SXM5 module.

OEM Example Platform Max H200 NVL GPUs
HPE ProLiant DL380a Gen11 / DL385 Gen11 Up to 4 (with NVLink bridge)
Dell PowerEdge XE9680 / R760xa Up to 4 (with NVLink bridge)
Supermicro AS-4125GS-TNRT2 / SYS-421GU-TNRT2 Up to 8
Lenovo ThinkSystem SR680a V3 Up to 4 (with NVLink bridge)
NVIDIA MGX Partner reference architecture Up to 8

Eton Technology supplies compatible HPE and Dell server platforms that support H200 NVL deployment, including the DL380a Gen11 and PowerEdge XE9680.

Browse Compatible Server Platforms

Deploying H200 NVL in UK Data Centres

Deploying H200 NVL GPUs in a UK data centre environment requires planning around power, cooling, and network infrastructure.

Power. Each H200 NVL has a configurable TDP up to 600W. A 4-GPU node draws up to 2.4kW for the GPUs alone, plus the host server. Standard UK data centre racks at 8-12kW per rack can hold 3-4 such nodes.

Cooling. The NVL is air-cooled, which simplifies deployment in existing facilities. For dense deployments (8 GPUs per node), supplemental fan kits or rear-door heat exchangers may be needed.

Networking. For multi-node scaling, NVIDIA recommends Spectrum-X Ethernet networking at 400GbE or higher. NVLink handles intra-node GPU communication, and NCCL manages inter-node communication over standard Ethernet or InfiniBand.

Lead time. H200 NVL availability has improved significantly since launch. Contact the Eton Technology team for current pricing and lead times.

H200 NVL for AI Inference and Training

The H200 NVL is designed primarily for AI inference, though it handles training workloads well within its memory and bandwidth envelope.

Workload Suitable config Notes
LLM inference (Llama 2 70B, GPT-3 class) 2-4 way NVLink Fits entirely in 2-4 GPU memory pool
LLM fine-tuning 4-way NVLink LoRA/QLoRA on single GPU, full fine-tune on 4
HPC simulation (MILC, NAMD, GROMACS) 2-4 way NVLink Benefits from 4.8TB/s memory bandwidth
RAG and embedding pipelines Single GPU Embedding models fit in 141GB easily
Batch inference (multi-tenant) Single GPU with MIG Partition into 7 instances for isolation
Video transcoding and processing Single GPU 7 NVDEC decoders handle heavy throughput

NVIDIA reports up to 1.7x faster LLM inference on the H200 NVL compared to the H100 NVL, driven by the larger 141GB memory pool and higher 4.8TB/s bandwidth. For UK enterprises running production inference, this translates to lower latency, higher throughput, and fewer GPUs needed to serve the same model.

Frequently Asked Questions

What does NVL stand for in NVIDIA GPUs?

NVL stands for NVLink. It indicates the GPU variant supports NVLink bridge connectivity, allowing multiple GPUs to share a high-speed interconnect for pooled memory and accelerated multi-GPU workloads.

What is the difference between H200 NVL and H200 HGX?

The H200 NVL is a PCIe dual-slot air-cooled card that fits into standard servers. The H200 HGX uses the SXM5 form factor and sits on a dedicated HGX baseboard. The GPU memory and compute performance are identical, but the NVL has a lower TDP (600W vs 700W) and includes a 5-year NVIDIA AI Enterprise subscription.

Can I use H200 NVL in an existing server?

Yes, provided the server has PCIe Gen5 slots with sufficient power delivery (600W per GPU) and physical clearance for dual-slot cards. HPE ProLiant, Dell PowerEdge, Supermicro, and Lenovo ThinkSystem platforms all support H200 NVL.

How many H200 NVL GPUs can I connect with NVLink?

The H200 NVL supports 2-way and 4-way NVLink bridge configurations. In a 4-way setup, four GPUs share a unified 564GB memory pool with 900GB/s bi-directional bandwidth per GPU.

Does the H200 NVL include NVIDIA AI Enterprise?

Yes. The H200 NVL comes with a 5-year NVIDIA AI Enterprise subscription included in the price. The H200 HGX (SXM5) requires AI Enterprise as a separate add-on purchase.

What is the lead time for H200 NVL in the UK?

Availability varies. Eton Technology can provide current lead time estimates and available configurations. Contact the team for up-to-date pricing and stock information.

Does Eton Technology supply H200 NVL GPUs to UK enterprises?

Yes. Eton Technology supplies H200 NVL GPUs, H200 HGX baseboards, and compatible HPE, Dell, and Lenovo server platforms to UK enterprises, research institutions, and government organisations.

Get in Touch

Eton Technology Solutions is a UK-based supplier of NVIDIA enterprise GPU infrastructure. The team supplies H200 NVL GPUs, H200 HGX baseboards, and compatible server platforms from HPE, Dell, Lenovo, and Supermicro.

Browse NVIDIA GPUs Shop H200 NVL

Phone: (+44) 333 999 7768 | Email: info@etontechnology.com

Cart 0

Your cart is currently empty.

Start Shopping
Call Us