Cisco is extending its Secure AI Factory with NVIDIA to liquid- and air-cooled Supermicro rack-scale systems from October, wrapping its own networking, validation and lifecycle services around them. A useful signal on how AI infrastructure is being packaged - and priced.
What was announced
- Cisco is expanding its Secure AI Factory with NVIDIA to rack-scale infrastructure built on liquid-cooled and air-cooled Supermicro systems, with availability starting in October through Cisco enterprise sales and channel.
- The design is anchored on NVIDIA Cloud Partner (NCP) Reference Architecture compliance plus Cisco Validated Infrastructure Services, Cisco AI networking and unified operations.
- Frontend fabric uses Cisco N9300 Series switches (Cisco Silicon One); backend fabric uses N9100 Series switches built on NVIDIA Spectrum-X silicon.
- Cisco maps the design from roughly 1,000 to more than 100,000 GPUs, across Vera Rubin NVL72 and GB300 NVL72 rack-scale systems, HGX Rubin NVL8 and HGX B300 NVL8, and MGX PCIe GPU platforms.
- An Enterprise Reference Architecture covers clusters under roughly 1,000 GPUs on N9300 switching; a Cloud Reference Architecture covers ~1,000 to 100,000+ GPUs.
- Cisco Cloud Control, the unified management console, is planned for calendar Q4 2026.
- Per ServeTheHome, physical hardware service on the Supermicro systems is expected to be performed by Supermicro.
The ETON view
This is the AI build market consolidating into pre-priced bundles, and it cuts both ways. For a genuine 1,000+ GPU greenfield cluster, a validated reference architecture with one throat to choke is worth paying for. But most enterprise and hosting AI projects we quote are 4 to 64 GPUs - an inference tier, a fine-tuning box, a departmental cluster - and at that scale a reference-architecture bundle buys you integration overhead and a lifecycle services wrapper you may already have in-house. Note also where the service boundary lands: the badge on the paperwork is not necessarily the engineer who replaces the GPU. Our view is to separate the decisions - take the reference architecture as a design pattern (Spectrum-X backend, separate frontend fabric, sane management network), then source the compute, networking and storage on availability and lead time. That is usually a materially lower capital number for the same topology, and it is how we build GPU platforms for hosting providers and AI companies today.
Related infrastructure
Category: AI & GPU · Vendor: Cisco, Supermicro, NVIDIA · Technology: GB300 NVL72, Vera Rubin NVL72, HGX B300 NVL8, MGX, Spectrum-X · Last verified: 02 Sep 2026 · ~2 min read
Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.
