What Is the NVIDIA H200 GPU?
The H200 is a data centre GPU built for generative AI, large language model inference, and high-performance computing. It shares the same GH100 die and Transformer Engine as the H100, so raw compute throughput is identical. The real upgrade is the memory subsystem: 141 GB of HBM3e running at 4.8 TB/s, compared to the H100's 80 GB at 3.35 TB/s.
That 76 per cent increase in memory capacity changes what's possible on a single GPU. Models with 70 billion parameters — like Llama 2 70B — can now run at FP16 precision on one H200 without model parallelism or quantisation. For UK research teams and enterprise AI deployments, that means simpler infrastructure and lower inference latency.
Key specs for the H200 SXM:
- GPU Memory: 141 GB HBM3e
- Memory Bandwidth: 4.8 TB/s
- FP8 Tensor Core Performance: 3,958 TFLOPS (with sparsity)
- FP16/BF16 Tensor Core Performance: 1,979 TFLOPS
- CUDA Cores: 16,896
- NVLink Bandwidth: 900 GB/s bidirectional per GPU
- Max TDP: 700W (configurable)
- Form Factor: SXM5 (requires HGX H200 or DGX H200 server platform)
- Multi-Instance GPU: Up to 7 MIG instances at 18 GB each
The H200 is also available in an NVL (PCIe) form factor, offering the same 141 GB memory and 4.8 TB/s bandwidth at 600W with air-cooled support for standard enterprise racks. The NVL includes a five-year NVIDIA AI Enterprise subscription with NIM microservices.

H200 vs H100: Key Differences
The H200 and H100 share the same compute architecture, so raw TFLOPS are identical. The differences are entirely in the memory subsystem — and that has a direct impact on real-world AI performance.
| H200 | H100 | |
|---|---|---|
| Memory | 141 GB HBM3e | 80 GB HBM3 |
| Bandwidth | 4.8 TB/s | 3.35 TB/s |
| Inference throughput | Up to 1.8x faster (memory-bound) | Baseline |
| Max TDP (SXM) | 700W | 700W |
| Cooling (NVL) | Air-cooled at 600W | Air-cooled at 600W |
For workloads that fit within 80 GB, the H100 remains a strong and more cost-effective option. For memory-bound inference on larger models, the H200 delivers meaningful gains that justify the premium. You can browse our NVIDIA GPU range to compare available options across both generations.
Why UK Enterprises Are Investing in H200 Infrastructure
AI infrastructure investment in the UK is moving fast. Microsoft has committed $30 billion to UK AI infrastructure through 2028, and NVIDIA is partnering with CoreWeave and Nscale to build AI factories across the country. The UK government's AI Growth Zones programme is accelerating data centre planning and power delivery to keep pace.
For UK enterprises, the H200 aligns well with several priorities:
Sovereign AI: Running models on UK-based infrastructure keeps sensitive data within national borders and reduces reliance on overseas cloud providers.
Energy efficiency: The H200 delivers more throughput per watt than the H100 for memory-bound workloads, which matters for organisations with sustainability targets.
Scalability: NVLink enables near-linear scaling across multi-GPU configurations, from 4-GPU HGX baseboards to large-scale deployments.
Enterprise support: The H200 NVL includes NVIDIA AI Enterprise with validated software stacks and direct NVIDIA support built in.
Deployment Considerations for UK Data Centres
Getting H200 infrastructure into a UK data centre requires some planning around power, cooling, and form factor.
Power and cooling
The H200 SXM runs at up to 700W per GPU and typically requires liquid cooling. If your facility isn't liquid-cooled, the H200 NVL is worth considering — it runs at 600W and is compatible with standard air-cooled racks, making it more accessible for existing UK data centre environments.
Server platforms
The H200 SXM requires an NVIDIA HGX H200 or DGX H200 platform in a 4-GPU or 8-GPU configuration. The H200 NVL fits into compatible HPE, Dell, and Supermicro platforms, which gives you more flexibility. Take a look at our server range to see what's available across both form factors.
Power infrastructure
An 8-GPU HGX H200 system can draw over 6 kW under load. It's worth auditing your existing power distribution and cooling headroom before committing to a deployment. The AI Growth Zones programme is designed to help data centre operators address exactly these challenges.
Lead times
Enterprise lead times for H200 GPUs typically run from three to eight weeks depending on configuration and volume. Get in touch and we can give you a current estimate.

Pricing and Procurement
Pricing varies by form factor, volume, and supplier. Current estimates place the H200 NVL (PCIe) at approximately £30,000 to £35,000 per GPU, with the SXM variant ranging from £38,000 to £45,000. Complete 8-GPU DGX H200 systems start at approximately £350,000.
For organisations looking to manage capital expenditure, we offer both new and refurbished server infrastructure to reduce the cost of building out AI capability.
When budgeting, factor in server platforms, networking (NVLink switches, InfiniBand or Ethernet fabric), cooling, and NVIDIA AI Enterprise licensing — included with the NVL, available as an add-on for SXM. We can provide a comprehensive quote covering all components.
Frequently Asked Questions
How much memory does the NVIDIA H200 have and why does it matter?
The H200 has 141 GB of HBM3e memory running at 4.8 TB/s — 76 per cent more capacity than the H100. That's enough to run 70-billion-parameter LLMs on a single GPU at FP16 precision, which simplifies deployment and cuts inference latency.
How does the H200 compare to the H100 for AI inference?
For compute-bound workloads, performance is identical. The H200's advantage is for memory-bound workloads — particularly LLM inference with long context windows. NVIDIA reports up to 1.8 times faster throughput for suitable workloads.
Is the H200 compatible with existing UK data centre infrastructure?
The SXM variant requires liquid cooling and HGX or DGX platforms. The NVL (PCIe) version is air-cooled and fits standard enterprise servers, making it the more accessible option for most existing UK facilities.
What is the typical lead time for H200 procurement in the UK?
Three to eight weeks for standard configurations, depending on volume and form factor. We can give you a current estimate based on what you need.
Can you provide a complete H200 server solution?
Yes — we supply NVIDIA GPUs, HPE and Dell servers, and refurbished infrastructure as an authorised UK reseller. Contact us for a tailored quote covering GPU, server platform, networking, and support.
Get in Touch
Whether you're evaluating the H200 for the first time or ready to spec out a full deployment, we're happy to help. Browse our NVIDIA infrastructure range or get in touch directly — we'll put together a quote tailored to your requirements.
Phone: +44 333 999 7768 | Email: info@etontechnology.com
