Skip to content
Eton Technology SolutionsEton Technology Solutions

Sizing a GPU server: what actually matters for inference versus training

September 02, 2026

Sizing a GPU server: what actually matters for inference versus training

Most AI hardware requirements we quote are far smaller than the headlines suggest. A working guide to specifying a GPU server for the job in front of you, without overbuying or hitting a wall in six months.

What to check

  • Decide first whether you are training, fine-tuning or serving. Serving is dominated by GPU memory capacity and bandwidth; training is dominated by interconnect and storage throughput.
  • Size GPU memory to the model plus its context and KV cache under concurrency, not to the model weights alone. Running out of memory at peak concurrency is the most common miss.
  • Check the PCIe topology of the chassis, not just the GPU count. How GPUs are attached to CPUs and NICs decides whether multi-GPU work scales.
  • For anything multi-node, specify the fabric at the same time as the compute. Network design is where small clusters most often underperform their spec sheet.
  • Provide fast local NVMe scratch. Data loading and checkpointing stall accelerators long before the GPU itself is the limit.
  • Confirm the power and cooling envelope per node against the rack you will install it in, including transient peaks rather than nameplate figures.
  • Check the software path: driver, container stack and framework support for the exact GPU model, especially on a newly launched part.
  • Leave a slot and power headroom for one more GPU than you think you need. Adding capacity later to a full chassis is the expensive route.

The ETON view

The gap between what AI infrastructure marketing describes and what a real project needs is enormous. A rack-scale reference architecture is the right answer for a frontier cluster and the wrong answer for a departmental inference tier, and the second is what most organisations are actually building. Two to eight GPUs, good local NVMe, sensible networking and a supported software stack covers a large share of the requirements we quote.

Buying that well is mostly about avoiding two failure modes. Overspecifying the accelerator while underspecifying memory, storage and network, which produces an expensive machine that idles. And specifying the newest part on the price list when a current-generation GPU with more available supply would ship in weeks instead of quarters, at a materially lower cost per unit of useful work.

We are vendor neutral, we hold stock, and we will happily talk you out of hardware you do not need. Send us the model, the concurrency you expect and the rack you have, and we will spec against that.

Related infrastructure

Category: AI & GPU · Vendor: ETON Technology · Technology: GPU memory, PCIe topology, NVLink, NVMe scratch · Last verified: 02 Sep 2026 · ~2 min read

Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.

Cart 0

Your cart is currently empty.

Start Shopping
Call Us