AMD used its IFA 2026 keynote to show a liquid-cooled workstation pairing a 96-core Threadripper PRO with Instinct MI350P accelerators, up to four of them for 576GB of HBM3E. No price, no date, no named OEM. We look at what the power and memory numbers mean if you are actually sizing local inference hardware.
What was announced
- Announced by Jack Huynh at AMD's IFA 2026 opening keynote on 4 September 2026 as a liquid-cooled tower workstation.
- The demo unit pairs a 96-core Threadripper PRO with two Instinct MI350P accelerators, with a stated path to four.
- AMD lists the Instinct MI350P at 144GB HBM3E, up to 4TB/s memory bandwidth, PCIe 5.0 x16, and 600W typical board power with a configurable 450W setting.
- Two cards give 288GB of accelerator memory; a four-card build reaches the 576GB figure quoted on stage.
- The only 96-core Threadripper PRO part is the 9995WX: Zen 5, 96 cores and 192 threads, up to 5.4 GHz boost, 384MB L3, 350W, 128 usable PCIe 5.0 lanes, AMD launch list price $11,699.
- AMD says the platform supports up to 2TB of eight-channel DDR5.
- AMD gave no pricing, no availability window, no named OEM partners and no benchmark data.
The ETON view
Strip out the keynote language and this is a specification exercise anyone can do today with parts that already exist. Two MI350P cards at 600W each plus a 350W CPU is 1,550W of silicon before drives, fans and conversion losses, which on a UK 13A socket leaves very little headroom and effectively forces a dedicated circuit or a four-card build straight into a rack. That is the real decision, not the branding: at four accelerators you are running datacentre-class power and heat in an office, and the sensible answer for most teams is a 4U GPU server in colo with the same cards in it.
The memory maths is worth checking against your own model rather than the trillion-parameter headline. A 4-bit quantised trillion-parameter model needs roughly 500GB just for weights, so only the full four-card 576GB configuration is in play, and those cards are PCIe attached rather than sharing one coherent pool. If your workload is fine-tuning or serving models in the 70B to 200B range, 288GB of HBM3E is already generous and the cost curve is much kinder.
On sourcing: there is no price, no date and no OEM, so nothing here is orderable. What is orderable now is MI350P-capable PCIe hosts and the current generation of HGX and PCIe GPU systems, and that is where we would put a Q4 budget. We build to the workload rather than the vendor, so if the requirement is local inference with data that cannot leave the building, we will quote both an AMD PCIe configuration and an NVIDIA one on lead time and total power draw, not on which keynote was louder.
Related infrastructure
Category: AI & GPU · Vendor: AMD, NVIDIA · Technology: Threadripper PRO 9995WX, Instinct MI350P, CDNA 4, HBM3E · Last verified: 06 Sep 2026 · ~2 min read
Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.
