Skip to content
Eton Technology SolutionsEton Technology Solutions
Data centre GPU comparison: L40S, H100, H200 and RTX PRO 6000 Blackwell specs side by side

Data centre GPU comparison: L40S, H100, H200 and RTX PRO 6000 Blackwell specs side by side

September 09, 2026

Data centre GPU comparison: L40S, H100, H200 and RTX PRO 6000 Blackwell specs side by side

Published specifications for the enterprise GPUs we are asked to quote most often, with the memory, bandwidth, power and interconnect differences that decide which one fits your server and your workload.

GPU choice is usually decided by three numbers: how much memory the model needs, how much power the chassis can deliver to a slot, and whether the workload needs GPU to GPU interconnect. Everything else is secondary.

Specification comparison

Vendor-published figures. PCIe and SXM variants of the same GPU differ, so confirm the variant in a quote.
GPU Memory Memory bandwidth Max board power NVLink Typical role
L40S 48 GB GDDR6 864 GB/s 350 W No Mixed AI, graphics, VDI, video, universal 2U workhorse
H100 (PCIe / SXM) 80 GB HBM2e / HBM3 Up to 3.35 TB/s (SXM) Up to 700 W (SXM) Yes Training and high-throughput inference
H200 141 GB HBM3e 4.8 TB/s Up to 700 W (SXM) Yes Large-model inference and training, memory-bound work
RTX PRO 6000 Blackwell Server Edition 96 GB GDDR7 Up to 1.6 TB/s 600 W No Inference density, consolidation, graphics plus AI
  • H200 offers 141 GB of HBM3e at 4.8 TB/s, roughly 1.8x the capacity and 1.4x the bandwidth of an 80 GB H100.
  • RTX PRO 6000 Blackwell Server Edition: 96 GB GDDR7, 512-bit interface, up to about 1.6 TB/s, 24,064 CUDA cores, 752 fifth-generation Tensor Cores, 4 PFLOPS FP4 and 2 PFLOPS FP8 tensor performance.
  • NVIDIA's reference architecture documentation notes an 8-GPU RTX PRO 6000 node reaches 768 GB of GDDR7 and up to 12.8 TB/s aggregate bandwidth, double the capacity and bandwidth of the equivalent L40S configuration.
  • L40S supports 3x NVENC and 3x NVDEC with AV1, is NEBS Level 3 ready, and does not support MIG.
  • Only the HBM parts (H100, H200) offer NVLink. PCIe-only cards scale by data parallelism, not by pooling memory.

What actually constrains the choice

  • Model memory footprint: a model that does not fit in one GPU's memory needs either NVLink-connected HBM parts or a smaller quantisation. Work out the footprint before shortlisting.
  • Slot power and cooling: 350 W (L40S), 600 W (RTX PRO 6000) and 700 W (SXM H100/H200) are three different mechanical and thermal problems. Many 2U servers are certified for a limited number of 300 to 400 W double-width cards.
  • Server certification: platform vendors publish supported GPU counts per chassis. The slot may be physically present and not thermally supported.
  • Precision support: Blackwell parts add FP4, which is what makes their inference throughput numbers so much higher than Ada or Hopper on the same power.
  • Licensing: virtualised GPU deployments need the relevant NVIDIA software licence, which is a real line item.

A sensible default per workload

  • VDI, rendering, transcode, light inference: L40S. Well supported, 350 W, fits most 2U servers.
  • Inference at density, or consolidating several L40S nodes: RTX PRO 6000 Blackwell Server Edition.
  • Training, fine-tuning, large-context inference: H100 or H200 in an HGX or SXM platform with NVLink.
  • Memory-bound inference where capacity per GPU decides throughput: H200.

The ETON view

Most GPU projects we quote are 4 to 64 GPUs, not thousand-GPU clusters, and at that scale the cost driver is the platform around the GPU rather than the GPU itself: power distribution, cooling, the backend fabric, and whether you need NVLink at all. A lot of teams specify HBM parts for workloads that would run comfortably on 96 GB of GDDR7 at a fraction of the capital cost and power.

Our advice is to start from the model and the throughput target, then work backwards to memory footprint, then to slot power, and only then choose the card. That order avoids the two expensive mistakes: buying training-class hardware for inference, and buying inference cards for a model that will not fit in them.

Availability is the other half. Lead times on HBM parts move constantly while L40S and PCIe Blackwell cards are far easier to source, and we hold and broker across the secondary market as well as the channel. Send us the workload and we will price the realistic options with honest lead times against each.

Related infrastructure

Category: AI & GPU · Vendor: NVIDIA · Technology: L40S, H100, H200, RTX PRO 6000 Blackwell Server Edition · Last verified: 09 Sep 2026 · ~3 min read

Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.

Cart 0

Your cart is currently empty.

Start Shopping
Call Us