Skip to content
AI Infrastructure

The Inference Reckoning: Custom Silicon, Sovereign Compute, and the End of Public-Cloud AI Economics

Google's Ironwood, AWS's Trainium3, and up to 1M TPUs for Anthropic — custom AI silicon is real and shipping. Nvidia still owns the merchant layer, sourced and ranged.

Published · How this was researched

Back to Global Tech Search

Market Horizon

2026 – 2027 (current hyperscaler capex cycle)

Target Sector

Cloud Infrastructure, AI Accelerators, Sovereign AI Deals

Anthropic TPU commitment

1M TPUs

>1GW of capacity coming online in 2026 (Google Cloud)

Nvidia accelerator share

81%

IDC estimate — other trackers range 75-87% (see chart)

MSFT CY2026 capex

~$190B

FY2026 Q3 earnings call guidance, Apr 29 2026

For a decade, "AI infrastructure" mostly meant one company's GPUs, rented by the hour, from whichever cloud happened to have capacity. That's no longer the whole picture. Google, Amazon, and Microsoft are now shipping accelerator chips they designed themselves, at a scale large enough that a frontier AI lab will build a meaningful chunk of its compute on top of them. The question this raises isn't whether custom silicon is real — it clearly is, with published specs and real deployments — but who actually ends up owning the infrastructure layer of AI, and whether "sovereignty" over compute is more than a rebranding of the same three or four companies.

Here's what I could verify, what's still single-sourced and should be read that way, and where the actual market-share numbers disagree with each other enough that citing just one would be misleading.

Google's Ironwood and the bet on TPUs at inference scale

Google's seventh-generation TPU, codenamed Ironwood, reached general availability in late 2025 after being announced at Cloud Next in April 2025 — and unlike earlier TPU generations, it's explicitly positioned as an inference chip rather than a training-first part. Per chip, Ironwood delivers 4,614 FP8 TFLOPS, 192GB of HBM3E memory, and 7.37 TB/s of memory bandwidth; wired together, a full pod scales to 9,216 chips delivering 42.5 exaflops. Google's own comparisons put Ironwood at 10x the peak performance of TPU v5p, more than 4x the per-chip performance of the prior generation (Trillium, TPU v6), roughly 2x the performance-per-watt of Trillium, and about 30x the performance of Google's first Cloud TPU. Trillium itself was no small step — 4.7x the peak performance per chip of TPU v5e, with 67% better performance per watt.

The number that actually moves this from "impressive spec sheet" to "structural shift in the market" is Anthropic's. In October 2025, Anthropic committed to accessing up to one million TPUs, with more than a gigawatt of that capacity coming online through 2026. That's not a pilot or a research collaboration — it's a frontier AI lab betting real production capacity on non-Nvidia silicon, at a scale that would have been unthinkable to describe casually even two years ago. Gemini 3, which launched November 18, 2025, is itself trained and served on TPUs, which tells you Google isn't just building this hardware for external customers; it's the same stack running its own frontier model.

AWS and Microsoft's answers

AWS's answer is Trainium3, launched December 2025 as the company's first 3-nanometer chip: 2.52 PFLOPS of FP8 compute, 144GB of HBM3e, and 4.9 TB/s of bandwidth. It sits on top of a fleet AWS has already built out — more than 500,000 Trainium2 chips deployed, described as the largest non-Nvidia AI accelerator cluster in production anywhere.

Microsoft's entry is Maia, and here I have to be careful about what's actually confirmed. A spec sheet reportedly describing "Maia 200" — TSMC 3nm process, 216GB of HBM3e, 7 TB/s bandwidth, roughly 750W power draw, dated January 2026 — has circulated via a single spec-aggregator source. Microsoft itself hasn't confirmed those numbers directly as far as I can find, so I'm presenting them as reported, not verified, and you should weight them accordingly. Meta's MTIA chips round out the fourth major hyperscaler entry into custom silicon, though with less public detail available.

Custom AI Silicon: Spec Comparison

Vendor-published specs; Maia 200 figures reported, unconfirmed by Microsoft

Ironwood (TPU v7)

4,614 TFLOPS FP8 · 192GB HBM

Trainium3

2,520 TFLOPS FP8 · 144GB HBM

Maia 200 (reported)

216GB HBM

Solid bar = FP8 TFLOPS/chip, outlined bar = HBM capacity. Maia 200 figures are single-source reported specs, unconfirmed by Microsoft. Vendor data.

Nvidia still owns the merchant layer

None of this means Nvidia's position has collapsed — it means it's narrowing, and even that narrowing depends on whose estimate you trust. Nvidia's share of the AI-accelerator market has been tracked at roughly 85–87% in 2024 by one research firm, around 81% by IDC, and projected toward roughly 75% for 2026 by another analyst group; AMD sits somewhere around 5–7%, Intel near 1%. I'm giving you that range deliberately rather than picking the flattest, cleanest-sounding number, because the methodology and the date matter more than the headline figure here.

The more durable point is structural, not statistical: not one of the custom hyperscaler chips — Ironwood, Trainium3, Maia — can be rented outside its own maker's cloud. If you want a TPU, you rent it from Google. If you want Trainium, you rent it from AWS. Nvidia GPUs, by contrast, are available across essentially every cloud provider and on-prem deployment that wants them, backed by CUDA, a software ecosystem competitors have spent a decade trying to dislodge without much success. Google's own Blackwell/GB200-class competition from Nvidia hit peak mass production in 2025, so this isn't a story of Nvidia falling behind on raw performance — it's a story of hyperscalers deciding that owning the chip, even a captive one, is worth the capital outlay it takes to stop paying Nvidia's margin on every unit.

Nvidia AI Accelerator Market Share

Presented as a range across trackers — IDC; Persistence Market Research; Silicon Analysts

  • Persistence Mkt Research (2024)85–87%
  • IDC81%
  • Silicon Analysts (2026 proj.)75%

AMD ~5–7%, Intel ~1% across the same trackers. Range shown deliberately — methodology and date change the figure. IDC; Persistence Market Research; Silicon Analysts.

The capital behind the bet

That capital outlay is enormous. Microsoft alone spent roughly $80 billion on AI infrastructure in fiscal year 2025 (ended June 2025), a figure Vice Chair and President Brad Smith cited directly in "The Golden Opportunity for American AI." On the company's fiscal 2026 Q3 earnings call (April 29, 2026), CFO Amy Hood guided to roughly $190 billion in calendar-2026 capex — including about $25 billion attributable to higher component pricing alone. Whatever fraction of that goes to custom silicon versus Nvidia hardware, the scale of spend confirms this isn't a hedge; it's the primary infrastructure bet of the decade for at least one hyperscaler, and the other two aren't spending meaningfully less. I've written before about the underlying inference cost curve this capital is chasing — that piece maps what a token costs to serve; this one is about who owns the metal serving it, which is a related but genuinely separate question.

Microsoft AI Capex Trajectory

Microsoft (Brad Smith blog; FY2026 Q3 earnings call, Apr 29 2026)

FY24 figure is an approximate prior-year baseline for trend context — Microsoft (Brad Smith blog; FY2026 Q3 earnings call, Apr 29 2026)

Sovereign compute: hyperscalers underwriting national AI stacks

The other axis reshaping this market is geopolitical rather than technical. Microsoft's $1.5 billion investment in the UAE's G42, announced April 16, 2024, grew into a $15.2 billion total UAE commitment through 2029, announced November 3, 2025 — a US hyperscaler directly underwriting a national AI compute stack for a Gulf state. Similar dynamics show up, with less primary confirmation, around Saudi Arabia's HUMAIN initiative and India's IndiaAI Mission GPU tender (reported at roughly $1.25 billion, or ₹10,371 crore) — I'm labeling that figure reported rather than verified, since I couldn't confirm it against a primary government filing. What's consistent across these deals is the pattern: "sovereign compute" in 2026 mostly means a country buying assurance and access from one of the same three or four US hyperscalers building this hardware, not building an independent chip supply chain of its own. That's worth naming plainly rather than letting the word "sovereign" do more rhetorical work than the underlying deal structure actually supports — a dynamic that connects directly to the broader map of where data centers and electricity capacity are actually being built globally right now.

The view from Kathmandu

Nepal rents essentially every FLOP it uses — there's no domestic hyperscaler, no sovereign chip program, no seat at any of these negotiating tables. Watching the G42 deal, or the IndiaAI tender, or Anthropic's TPU commitment, isn't abstract from here; it's watching, in real time, which countries get invited into the infrastructure layer of the next decade of computing and which remain purely on the renting side of it. The custom-silicon story reads, in US tech coverage, as a story about hyperscaler competition. From here it reads as a story about which governments get a phone call before the compute gets allocated — and which don't.

Sources

  • Google Cloud Blog; blog.google — "Ironwood: first TPU for the age of inference," specs and generational comparisons, 2025
  • Google Cloud Blog — Anthropic TPU commitment (up to 1M TPUs, >1GW in 2026), October 2025
  • AWS/Introl — Trainium3 launch and specifications, December 2025; Hashrate Index — Trainium2 fleet size (500,000+ chips)
  • Microsoft — "The Golden Opportunity for American AI" (Brad Smith), FY25 ~$80B AI capex; Microsoft FY2026 Q3 earnings call transcript (April 29, 2026), ~$190B calendar-2026 guidance
  • news.microsoft.com — Microsoft–G42 $1.5B investment, April 16, 2024; blogs.microsoft.com — $15.2B UAE commitment, November 3, 2025
  • IDC; Persistence Market Research; Silicon Analysts — Nvidia AI-accelerator market share estimates, 2024–2026 (cited as a range, not a single figure)
  • TrendForce — Nvidia Blackwell/GB200 mass production status, 2025
  • Spheron (spec aggregator) — Microsoft Maia 200 datasheet, reported January 2026, unconfirmed directly by Microsoft
  • Google — Gemini 3 launch announcement, November 18, 2025

Written by Abhishek Kushwaha, founder and writer at Global Tech Search, based in Kathmandu, Nepal.

Advantages

  • Anthropic's commitment to access up to one million Google TPUs is the clearest proof yet that a frontier lab will build at real scale on non-Nvidia silicon
  • AWS Trainium3 is a real, shipped 3nm part backing a 500,000+ chip Trainium2 fleet already in production, not a roadmap slide
  • Microsoft's ~$190B calendar-2026 capex guidance shows the capital markets treating this shift as durable, not speculative

× Challenges

  • None of the hyperscaler ASICs are rentable outside their owners' own clouds, so lock-in isn't solved here — it's relocated from Nvidia to Google, Amazon, or Microsoft
  • Nvidia still holds a clear majority of the merchant accelerator market by every tracker's estimate, even as the exact share varies by source and quarter
  • Microsoft's Maia 200 spec sheet is single-sourced and unconfirmed directly by Microsoft — treated here as reported, not verified

Risk Assessment

"Sovereign compute" and "custom silicon" both still route through three or four US hyperscalers — owning the chip design doesn't yet mean owning the fabrication, the packaging, or the CUDA-equivalent software layer that makes a chip usable at scale.

Abhishek Kushwaha

Written by Abhishek Kushwaha

Full-stack software engineer in Kathmandu, Nepal — six years shipping production Django and Next.js systems, most recently at Pinakin Technologies & Research Center. Writes Global Tech Search on what the AI buildout costs in power, water and silicon, measuring it first-hand where he can. More about the author →