Published specifications for the enterprise GPUs we are asked to quote most often, with the memory, bandwidth, power and interconnect differences that decide which one fits your server and your workload.
GPU choice is usually decided by three numbers: how much memory the model needs, how much power the chassis can deliver to a slot, and whether the workload needs GPU to GPU interconnect. Everything else is secondary.
Specification comparison
| GPU | Memory | Memory bandwidth | Max board power | NVLink | Typical role |
|---|---|---|---|---|---|
| L40S | 48 GB GDDR6 | 864 GB/s | 350 W | No | Mixed AI, graphics, VDI, video, universal 2U workhorse |
| H100 (PCIe / SXM) | 80 GB HBM2e / HBM3 | Up to 3.35 TB/s (SXM) | Up to 700 W (SXM) | Yes | Training and high-throughput inference |
| H200 | 141 GB HBM3e | 4.8 TB/s | Up to 700 W (SXM) | Yes | Large-model inference and training, memory-bound work |
| RTX PRO 6000 Blackwell Server Edition | 96 GB GDDR7 | Up to 1.6 TB/s | 600 W | No | Inference density, consolidation, graphics plus AI |
- H200 offers 141 GB of HBM3e at 4.8 TB/s, roughly 1.8x the capacity and 1.4x the bandwidth of an 80 GB H100.
- RTX PRO 6000 Blackwell Server Edition: 96 GB GDDR7, 512-bit interface, up to about 1.6 TB/s, 24,064 CUDA cores, 752 fifth-generation Tensor Cores, 4 PFLOPS FP4 and 2 PFLOPS FP8 tensor performance.
- NVIDIA's reference architecture documentation notes an 8-GPU RTX PRO 6000 node reaches 768 GB of GDDR7 and up to 12.8 TB/s aggregate bandwidth, double the capacity and bandwidth of the equivalent L40S configuration.
- L40S supports 3x NVENC and 3x NVDEC with AV1, is NEBS Level 3 ready, and does not support MIG.
- Only the HBM parts (H100, H200) offer NVLink. PCIe-only cards scale by data parallelism, not by pooling memory.
What actually constrains the choice
- Model memory footprint: a model that does not fit in one GPU's memory needs either NVLink-connected HBM parts or a smaller quantisation. Work out the footprint before shortlisting.
- Slot power and cooling: 350 W (L40S), 600 W (RTX PRO 6000) and 700 W (SXM H100/H200) are three different mechanical and thermal problems. Many 2U servers are certified for a limited number of 300 to 400 W double-width cards.
- Server certification: platform vendors publish supported GPU counts per chassis. The slot may be physically present and not thermally supported.
- Precision support: Blackwell parts add FP4, which is what makes their inference throughput numbers so much higher than Ada or Hopper on the same power.
- Licensing: virtualised GPU deployments need the relevant NVIDIA software licence, which is a real line item.
A sensible default per workload
- VDI, rendering, transcode, light inference: L40S. Well supported, 350 W, fits most 2U servers.
- Inference at density, or consolidating several L40S nodes: RTX PRO 6000 Blackwell Server Edition.
- Training, fine-tuning, large-context inference: H100 or H200 in an HGX or SXM platform with NVLink.
- Memory-bound inference where capacity per GPU decides throughput: H200.
The ETON view
Most GPU projects we quote are 4 to 64 GPUs, not thousand-GPU clusters, and at that scale the cost driver is the platform around the GPU rather than the GPU itself: power distribution, cooling, the backend fabric, and whether you need NVLink at all. A lot of teams specify HBM parts for workloads that would run comfortably on 96 GB of GDDR7 at a fraction of the capital cost and power.
Our advice is to start from the model and the throughput target, then work backwards to memory footprint, then to slot power, and only then choose the card. That order avoids the two expensive mistakes: buying training-class hardware for inference, and buying inference cards for a model that will not fit in them.
Availability is the other half. Lead times on HBM parts move constantly while L40S and PCIe Blackwell cards are far easier to source, and we hold and broker across the secondary market as well as the channel. Send us the workload and we will price the realistic options with honest lead times against each.
Related infrastructure
Category: AI & GPU · Vendor: NVIDIA · Technology: L40S, H100, H200, RTX PRO 6000 Blackwell Server Edition · Last verified: 09 Sep 2026 · ~3 min read
Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.
