For a decade, "AI infrastructure" mostly meant one company's GPUs, rented by the hour, from whichever cloud happened to have capacity. That's no longer the whole picture. Google, Amazon, and Microsoft are now shipping accelerator chips they designed themselves, at a scale large enough that a frontier AI lab will build a meaningful chunk of its compute on top of them. The question this raises isn't whether custom silicon is real — it clearly is, with published specs and real deployments — but who actually ends up owning the infrastructure layer of AI, and whether "sovereignty" over compute is more than a rebranding of the same three or four companies.
Here's what I could verify, what's still single-sourced and should be read that way, and where the actual market-share numbers disagree with each other enough that citing just one would be misleading.
Google's Ironwood and the bet on TPUs at inference scale
Google's seventh-generation TPU, codenamed Ironwood, reached general availability in late 2025 after being announced at Cloud Next in April 2025 — and unlike earlier TPU generations, it's explicitly positioned as an inference chip rather than a training-first part. Per chip, Ironwood delivers 4,614 FP8 TFLOPS, 192GB of HBM3E memory, and 7.37 TB/s of memory bandwidth; wired together, a full pod scales to 9,216 chips delivering 42.5 exaflops. Google's own comparisons put Ironwood at 10x the peak performance of TPU v5p, more than 4x the per-chip performance of the prior generation (Trillium, TPU v6), roughly 2x the performance-per-watt of Trillium, and about 30x the performance of Google's first Cloud TPU. Trillium itself was no small step — 4.7x the peak performance per chip of TPU v5e, with 67% better performance per watt.
The number that actually moves this from "impressive spec sheet" to "structural shift in the market" is Anthropic's. In October 2025, Anthropic committed to accessing up to one million TPUs, with more than a gigawatt of that capacity coming online through 2026. That's not a pilot or a research collaboration — it's a frontier AI lab betting real production capacity on non-Nvidia silicon, at a scale that would have been unthinkable to describe casually even two years ago. Gemini 3, which launched November 18, 2025, is itself trained and served on TPUs, which tells you Google isn't just building this hardware for external customers; it's the same stack running its own frontier model.
AWS and Microsoft's answers
AWS's answer is Trainium3, launched December 2025 as the company's first 3-nanometer chip: 2.52 PFLOPS of FP8 compute, 144GB of HBM3e, and 4.9 TB/s of bandwidth. It sits on top of a fleet AWS has already built out — more than 500,000 Trainium2 chips deployed, described as the largest non-Nvidia AI accelerator cluster in production anywhere.
Microsoft's entry is Maia, and here I have to be careful about what's actually confirmed. A spec sheet reportedly describing "Maia 200" — TSMC 3nm process, 216GB of HBM3e, 7 TB/s bandwidth, roughly 750W power draw, dated January 2026 — has circulated via a single spec-aggregator source. Microsoft itself hasn't confirmed those numbers directly as far as I can find, so I'm presenting them as reported, not verified, and you should weight them accordingly. Meta's MTIA chips round out the fourth major hyperscaler entry into custom silicon, though with less public detail available.
Custom AI Silicon: Spec Comparison
Vendor-published specs; Maia 200 figures reported, unconfirmed by Microsoft
Ironwood (TPU v7)
4,614 TFLOPS FP8 · 192GB HBM
Trainium3
2,520 TFLOPS FP8 · 144GB HBM
Maia 200 (reported)
216GB HBM
Solid bar = FP8 TFLOPS/chip, outlined bar = HBM capacity. Maia 200 figures are single-source reported specs, unconfirmed by Microsoft. Vendor data.
Nvidia still owns the merchant layer
None of this means Nvidia's position has collapsed — it means it's narrowing, and even that narrowing depends on whose estimate you trust. Nvidia's share of the AI-accelerator market has been tracked at roughly 85–87% in 2024 by one research firm, around 81% by IDC, and projected toward roughly 75% for 2026 by another analyst group; AMD sits somewhere around 5–7%, Intel near 1%. I'm giving you that range deliberately rather than picking the flattest, cleanest-sounding number, because the methodology and the date matter more than the headline figure here.
The more durable point is structural, not statistical: not one of the custom hyperscaler chips — Ironwood, Trainium3, Maia — can be rented outside its own maker's cloud. If you want a TPU, you rent it from Google. If you want Trainium, you rent it from AWS. Nvidia GPUs, by contrast, are available across essentially every cloud provider and on-prem deployment that wants them, backed by CUDA, a software ecosystem competitors have spent a decade trying to dislodge without much success. Google's own Blackwell/GB200-class competition from Nvidia hit peak mass production in 2025, so this isn't a story of Nvidia falling behind on raw performance — it's a story of hyperscalers deciding that owning the chip, even a captive one, is worth the capital outlay it takes to stop paying Nvidia's margin on every unit.
Nvidia AI Accelerator Market Share
Presented as a range across trackers — IDC; Persistence Market Research; Silicon Analysts
- Persistence Mkt Research (2024)85–87%
- IDC81%
- Silicon Analysts (2026 proj.)75%
AMD ~5–7%, Intel ~1% across the same trackers. Range shown deliberately — methodology and date change the figure. IDC; Persistence Market Research; Silicon Analysts.
The capital behind the bet
That capital outlay is enormous. Microsoft alone spent roughly $80 billion on AI infrastructure in fiscal year 2025 (ended June 2025), a figure Vice Chair and President Brad Smith cited directly in "The Golden Opportunity for American AI." On the company's fiscal 2026 Q3 earnings call (April 29, 2026), CFO Amy Hood guided to roughly $190 billion in calendar-2026 capex — including about $25 billion attributable to higher component pricing alone. Whatever fraction of that goes to custom silicon versus Nvidia hardware, the scale of spend confirms this isn't a hedge; it's the primary infrastructure bet of the decade for at least one hyperscaler, and the other two aren't spending meaningfully less. I've written before about the underlying inference cost curve this capital is chasing — that piece maps what a token costs to serve; this one is about who owns the metal serving it, which is a related but genuinely separate question.
Microsoft AI Capex Trajectory
Microsoft (Brad Smith blog; FY2026 Q3 earnings call, Apr 29 2026)
FY24 figure is an approximate prior-year baseline for trend context — Microsoft (Brad Smith blog; FY2026 Q3 earnings call, Apr 29 2026)
Sovereign compute: hyperscalers underwriting national AI stacks
The other axis reshaping this market is geopolitical rather than technical. Microsoft's $1.5 billion investment in the UAE's G42, announced April 16, 2024, grew into a $15.2 billion total UAE commitment through 2029, announced November 3, 2025 — a US hyperscaler directly underwriting a national AI compute stack for a Gulf state. Similar dynamics show up, with less primary confirmation, around Saudi Arabia's HUMAIN initiative and India's IndiaAI Mission GPU tender (reported at roughly $1.25 billion, or ₹10,371 crore) — I'm labeling that figure reported rather than verified, since I couldn't confirm it against a primary government filing. What's consistent across these deals is the pattern: "sovereign compute" in 2026 mostly means a country buying assurance and access from one of the same three or four US hyperscalers building this hardware, not building an independent chip supply chain of its own. That's worth naming plainly rather than letting the word "sovereign" do more rhetorical work than the underlying deal structure actually supports — a dynamic that connects directly to the broader map of where data centers and electricity capacity are actually being built globally right now.
The view from Kathmandu
Nepal rents essentially every FLOP it uses — there's no domestic hyperscaler, no sovereign chip program, no seat at any of these negotiating tables. Watching the G42 deal, or the IndiaAI tender, or Anthropic's TPU commitment, isn't abstract from here; it's watching, in real time, which countries get invited into the infrastructure layer of the next decade of computing and which remain purely on the renting side of it. The custom-silicon story reads, in US tech coverage, as a story about hyperscaler competition. From here it reads as a story about which governments get a phone call before the compute gets allocated — and which don't.
Sources
- Google Cloud Blog; blog.google — "Ironwood: first TPU for the age of inference," specs and generational comparisons, 2025
- Google Cloud Blog — Anthropic TPU commitment (up to 1M TPUs, >1GW in 2026), October 2025
- AWS/Introl — Trainium3 launch and specifications, December 2025; Hashrate Index — Trainium2 fleet size (500,000+ chips)
- Microsoft — "The Golden Opportunity for American AI" (Brad Smith), FY25 ~$80B AI capex; Microsoft FY2026 Q3 earnings call transcript (April 29, 2026), ~$190B calendar-2026 guidance
- news.microsoft.com — Microsoft–G42 $1.5B investment, April 16, 2024; blogs.microsoft.com — $15.2B UAE commitment, November 3, 2025
- IDC; Persistence Market Research; Silicon Analysts — Nvidia AI-accelerator market share estimates, 2024–2026 (cited as a range, not a single figure)
- TrendForce — Nvidia Blackwell/GB200 mass production status, 2025
- Spheron (spec aggregator) — Microsoft Maia 200 datasheet, reported January 2026, unconfirmed directly by Microsoft
- Google — Gemini 3 launch announcement, November 18, 2025
Written by Abhishek Kushwaha, founder and writer at Global Tech Search, based in Kathmandu, Nepal.
