Skip to content
Eton Technology SolutionsEton Technology Solutions

Qualcomm and AWS Commit to Multi-Generation Inference Silicon: What Custom Accelerators Mean for Everyone Else's Buying Plans

September 11, 2026

Qualcomm and AWS Commit to Multi-Generation Inference Silicon: What Custom Accelerators Mean for Everyone Else's Buying Plans

Qualcomm and Amazon have agreed a multi-generation collaboration on custom AI inference silicon and optical connectivity up to 1.6T for AWS data centres. The hyperscalers keep building their own inference stack, which shapes what is left on the open market for everyone else.

What was announced

  • Announced 8 September 2026: Qualcomm Technologies and Amazon will collaborate across multiple generations of customised silicon for large-scale AI data centres, with AI inference as the focus.
  • The two companies are also working on optical connectivity solutions extending up to 1.6T, drawing on Qualcomm SerDes and optical DSP technology.
  • Qualcomm plans to deepen its use of AWS AI infrastructure, including Amazon Bedrock, for electronic design automation workloads, aiming to shorten chip design cycles.
  • Qualcomm named no specific parts, delivery dates or financial terms in the announcement.

The ETON view

Every hyperscaler custom inference programme has the same effect on the rest of the market. Volume that would once have gone through the general purpose GPU channel now goes into silicon nobody outside AWS can buy, and the parts that do reach the channel get priced against that scarcity. If you are a hosting provider, a cloud platform or an enterprise sourcing your own inference capacity, the practical lesson is that you are not competing with AWS for the same chips, you are competing for the leftovers of a supply chain tuned to hyperscale orders.

That argues for two things. First, size inference on what you can actually buy and get racked this quarter, not on a roadmap part. Current generation accelerators with real UK stock and a known lead time will beat a next generation part that lands in nine months. Second, do not let the compute story hide the network. The interesting half of this deal is 1.6T optical, and the same bottleneck applies at a smaller scale: an inference cluster that is short on fabric bandwidth wastes the GPUs you paid for.

ETON works the open market rather than one vendor roadmap. We source current and previous generation GPU platforms, the switching and optics to go with them, and we quote on availability and lead time up front. Where a previous generation part gets an inference job done at a fraction of the cost per token, we will say so rather than sell you the newest badge.

Related infrastructure

Category: AI & GPU · Vendor: Qualcomm, Amazon, NVIDIA · Technology: AI inference accelerators, 1.6T optical interconnect, SerDes · Last verified: 10 Sep 2026 · ~2 min read

Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.

Cart 0

Your cart is currently empty.

Start Shopping
Call Us