Skip to content
Eton Technology SolutionsEton Technology Solutions
GPU server platforms compared: which rack servers take which GPUs, and how many

GPU server platforms compared: which rack servers take which GPUs, and how many

September 09, 2026

GPU server platforms compared: which rack servers take which GPUs, and how many

A guide to the server side of a GPU build: accelerator-optimised platforms versus general-purpose rack servers, how many double-width cards each class really supports, and the power, cooling and slot facts that decide the answer.

Choosing the GPU is the easy half. The platform decides how many you can fit, how hard you can run them, and whether they can talk to each other.

Three classes of platform

  • General-purpose 2U rack servers (DL380 Gen11, PowerEdge R760, Supermicro Ultra): typically two to three double-width cards, often power-capped, best for one or two inference or VDI GPUs alongside normal workloads.
  • Accelerator-optimised rack servers (HPE ProLiant Compute DL380a Gen12, Dell PowerEdge R760xa, Supermicro GPU SuperServers): purpose-built airflow and power for four to eight double-width PCIe GPUs. HPE's DL380a Gen12 CTO configuration is documented as 8 double-wide or 16 single-wide.
  • HGX and SXM systems (8-GPU HGX H100/H200 baseboards, rack-scale NVL platforms): NVLink-connected HBM GPUs, liquid or high-volume air cooling, the only route to pooled GPU memory.

What limits GPU count

  • Slot power: a 350 W L40S, a 600 W RTX PRO 6000 and a 700 W SXM part are three different design problems. Total board power often exceeds the CPU and drive power combined.
  • Airflow and inlet temperature: vendors publish thermal restrictions, and a supported GPU count usually assumes a maximum ambient. Above it, the count drops.
  • PSU capacity and PDU feed: eight 600 W cards plus platform overhead needs redundant supplies sized accordingly and a rack feed that can take it.
  • PCIe topology: how the slots are wired to CPUs and switches decides host-to-GPU bandwidth and whether peer-to-peer transfers are efficient.
  • Physical clearance and auxiliary power cabling: the connector, not the slot, is often the blocker on used platforms.

PCIe or NVLink

  • PCIe cards scale by data parallelism: each GPU holds its own copy of the model. Cheap, flexible, easy to source.
  • NVLink-connected SXM GPUs pool memory and bandwidth, which is what training and very large-context inference need.
  • NVIDIA's reference architecture documentation puts an 8-GPU RTX PRO 6000 node at 768 GB GDDR7 and up to 12.8 TB/s aggregate bandwidth, double the L40S equivalent, without NVLink. For many inference estates that is enough.
  • Decide this from the model footprint. If a single GPU holds the model, PCIe is almost always the better commercial answer.

A build checklist

  • Confirm the exact GPU SKU appears on the platform's supported option list for the number of cards you want.
  • Confirm rack power and cooling per rack unit with the colo or facilities team before ordering.
  • Size the network for the workload: a separate backend fabric for multi-node training, and a sensible frontend and management network.
  • Plan storage bandwidth. A GPU node starved of data is an expensive idle asset.
  • Check the software entitlement (drivers, virtualisation licences, container stack) as a line item, not an afterthought.

The ETON view

The mistake we see most often on GPU builds is buying the cards first and then discovering the platform, the rack or the hall cannot support them. Power and airflow are the real constraints in UK colo, where a lot of space is still provisioned around 6 to 12 kW per rack, and a fully populated accelerator node can consume a large part of that on its own.

We would rather design from the constraint inwards: what the hall gives you per rack, then the platform that fits it, then the GPU count, then the card. That order routinely produces a cheaper build that actually runs, and it often shows that two mid-power inference nodes beat one flagship node on both throughput and resilience.

We are vendor neutral across HPE, Dell, Supermicro, GIGABYTE and ASUS platforms and we source GPUs through the channel and the secondary market, so we can price several topologies for the same workload with honest lead times against each. Send us the model, the throughput target and your rack power and we will do exactly that.

Related infrastructure

Category: AI & GPU · Vendor: HPE, Dell, Supermicro, NVIDIA · Technology: DL380a Gen12, PowerEdge R760xa, Supermicro GPU systems, HGX vs PCIe · Last verified: 09 Sep 2026 · ~3 min read

Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.

Cart 0

Your cart is currently empty.

Start Shopping
Call Us