Process node
2nm
First high-volume phone SoC, Apple's claim
Neural Engine
32 cores
2x compute of predecessor, Apple's claim
Sustained perf
+40%
vs A19 Pro, Apple's claim
Energy per token
—
Not disclosed by anyone
Apple announced the iPhone 18 Pro, the iPhone 18 Pro Max, and its first foldable on September 9, 2026. Preorders open September 12; the Pro models ship September 18 at $1,199 and $1,299.
This site does not cover phone launches. It covers the physical footprint of AI compute — power, wafers, packaging, heat — and it covers this one because the A20 Pro is the first leading-edge silicon of 2026 to reach volume, it consumed a large share of the industry's scarcest manufacturing capacity to get there, and Apple used its keynote to make a specific, checkable claim about running language models locally. That claim belongs to this beat. The camera does not.
What Apple actually confirmed
Apple's own figures, from the keynote, none of them independently verified on the date of publication:
The A20 Pro is built on a 2nm process, which Apple describes as the first in a high-volume smartphone. The CPU is six cores — two new "super cores" Apple calls desktop class and the fastest in any smartphone, up to 20% faster, plus four efficiency cores. The GPU is seven cores, up to 40% faster at graphics. The Neural Engine is 32 cores with twice the compute of its predecessor, and its neural accelerators deliver twice the FP8 performance of the previous generation. Apple quoted sustained performance twice: 40% more than the A19 Pro, and twice that of the A18 Pro.
Two packaging details matter more than the core counts. Apple said the chip uses custom packaging "inspired by our M-series chips" that connects silicon and memory directly to the vapor chamber. And it said the A20 Pro carries the widest memory interface ever shipped in an iPhone, which 9to5Mac reported as roughly a 50% increase in memory bandwidth.
Apple's framing was explicit: "the ultimate chip for running advanced on-device models," with a redesigned Neural Engine supporting 8-bit floating point math, capable of running large language models locally.
The two specifications that are not marketing
Most of that list is ordinary generational improvement. Two items are not, and they are the two worth a reader's attention.
FP8 is a data center numeric format. Eight-bit floating point exists because transformer inference at scale needed a representation cheaper than FP16 without the accuracy collapse of naive integer quantisation. It is the format of H100-class accelerators and modern inference servers. Doubling FP8 throughput in a phone NPU is not a general "AI features" investment — it is a targeted bet on transformer inference specifically, and it tells you what Apple expects the silicon to spend its time doing. This site's quantisation tradeoff measurements cover what that numeric choice actually costs in output quality on Apple silicon, which is the half of the FP8 story no keynote covers.
Leading with sustained performance is an admission. Apple quoted sustained performance against two prior generations and bonded the memory directly to the vapor chamber. Peak throughput sells benchmarks; sustained throughput is what a minutes-long generation actually experiences, and the gap between them is thermal. By putting sustained numbers in the headline and redesigning the thermal path to support them, Apple is conceding that the binding constraint on phone inference is heat rejection, not arithmetic.
That is the same constraint, at a different scale, that governs the buildings this site usually writes about. It is also what this site's own sustained throughput runs on an M1 found: local inference on Apple silicon is a thermally limited workload long before it is a compute-limited one. The A20 Pro's most interesting engineering is a cooling story.
The wafers this chip came out of
A phone SoC and an AI accelerator are not in different industries. In 2026 they are in the same queue.
Leading-edge wafer capacity is the tightest input in the compute supply chain, and it is not segmented by end market. DigiTimes reported in August 2025 that Apple had ordered almost half of TSMC's initial 2nm capacity for the iPhone 18 generation, against an N2 ramp of roughly 45,000-50,000 wafers per month at the end of 2025 scaling past 100,000 per month during 2026.
That figure deserves care. It is single-sourced trade-press reporting from over a year ago, it describes initial capacity rather than the current or full-year split, and neither Apple nor TSMC has confirmed it. Treat it as directionally informative, not as a settled number — the same standard this site applied to the analyst wafer estimates in its CoWoS packaging bottleneck coverage, where TSMC's CEO named the constraint on the record but the company declined to attach wafer figures to it.
With that caveat, the structural point holds and is not seriously disputed: a very large fraction of the first 2nm output went into consumer devices. Every wafer that becomes an A20 Pro is a wafer that did not become something else, in a year when TSMC's own leadership has said capacity is limiting its customers' growth. The AI buildout and the iPhone are competing for the same physical output of the same small number of fabs — a dependency this site has traced before in Sovereign Compute, Splintered Silicon.
The useful inversion: a consumer product cycle is now a variable in AI infrastructure supply. Not a metaphor for one.
Innovation, increment, and experiment
Sorting the launch on engineering grounds rather than launch-day enthusiasm.
Structural — the 2nm node and the thermal path. A full node transition plus memory bonded to the vapor chamber changes the performance envelope rather than the feature list, and it is the part competitors cannot answer inside a product cycle. The sustained-performance figures are the claim to test.
Structural — FP8 throughput and memory bandwidth. Doubled FP8 and the widest memory interface Apple has shipped in a phone are the two specifications that decide whether a local model runs usefully or merely runs. Memory bandwidth in particular is the ceiling this site has measured directly on Apple silicon: token generation is bandwidth-bound long before it is FLOP-bound, which the inference memory bandwidth runs show in detail.
Increment — variable aperture, the smaller Dynamic Island, the C2 modem. Real engineering, genuinely useful, and none of it changes what the device is capable of in kind. The C2 modem matters commercially, as continued substitution away from a purchased component, but that is a supply-chain fact rather than a capability.
Experiment — the foldable, at two thousand dollars. Announced September 9, reported by MacRumors as starting at $2,000 and not launching until October at the earliest, under a name that originates with Bloomberg's reporting rather than Apple's own materials. A constrained first-generation form factor at that price is a market test, and should be read as one.
Experiment — Apple Reference Image. A provenance signal for verifying a photograph is not AI-generated. Interesting, and unproven: provenance systems fail on adoption and adversarial pressure, not on the quality of the initial implementation.
The number nobody published
Here is the gap this post exists to name.
Apple made a claim about running language models on the device. To evaluate it, you need to know what a token costs in joules. Not compute delivered, not a ratio against last year's chip — energy per token. That number decides the entire question the on-device AI pitch rests on, because "runs locally" is not the same as "costs less." Local inference does not make the energy disappear. It moves it from a data center's power contract to a battery the user recharges, and whether that trade is favourable is an empirical question with a real answer.
Nobody published it yesterday. Apple did not, and phone reviewers will not, and the reason is not that anyone is hiding something.
This site tried to measure exactly that quantity on Apple silicon and failed, in ways it published in full. The macOS battery gauge holds each value for 51-59 seconds while the generations being measured last 2.8 to 5.4 seconds. Where a window long enough to beat that refresh rate existed, two independent derivations of the same energy — integrating instantaneous power, and reading the battery pack's own coulomb counter — disagreed by 2.05x over the same 21-minute run. The analysis script emitted a marginal energy figure of exactly 0.00 J/token for two models, which reads like a clean result and is actually an instrument that never resolved the load.
That was on a Mac, with a shell, with unprivileged access to power telemetry that iOS does not expose at all. If per-query energy is not measurable on an M1 without root, it is not going to be measured on an iPhone by a reviewer with a stopwatch and a battery percentage.
So the honest position on September 10, 2026 is this: Apple has shipped hardware that is very likely much better at local inference, on specifications that are the right specifications to care about, and the claim that this makes AI cheaper or greener is unevidenced in both directions. This site's local versus cloud crossover work covers where that boundary sits on measurable axes — latency, throughput, cost — and energy is precisely the axis that stayed out of reach.
What would settle it
Three things, in rough order of tractability:
A tokens-per-second figure for a named model at a named quantisation on retail A20 Pro hardware, sustained over a run long enough to reach thermal equilibrium rather than a first-30-seconds burst. Any competent reviewer with a shipping device can produce this after September 18, and it would be the first real datapoint about the chip's inference behaviour.
A power figure taken at the wall during that sustained run — charge the device from empty at a known state while looping generation, and difference against an idle baseline. Crude, and far better than nothing, because it sidesteps the on-device telemetry that made the M1 measurement fail.
Disclosure from Apple of a wattage envelope for the Neural Engine under sustained load. Unlikely, and it is the number that would end the argument.
Until one of those exists, "the ultimate chip for running advanced on-device models" is a statement about compute delivered per second. It is not a statement about what the compute costs, and the two are routinely confused — including, frequently, by people who know better.
Sources
- Apple keynote claims for the A20 Pro via MacRumors and 9to5Mac, September 9, 2026 — every core count, ratio, and packaging detail attributed to Apple in this post comes from that day's coverage of Apple's own presentation, and none of it is independently verified
- DigiTimes via MacRumors, August 28, 2025 — the "almost half of initial 2nm capacity" allocation claim and the N2 wafer ramp figures, labeled throughout as single-sourced trade-press reporting that neither Apple nor TSMC has confirmed
- MacRumors, September 9, 2026 — foldable pricing and availability; the product name originates with Bloomberg's Mark Gurman rather than with Apple marketing materials
- This site's own committed measurement data — data/energy-per-token/ for the failed per-query energy measurement and its 2.05x reconciliation failure, and the sustained-throughput and memory-bandwidth run logs cited via the linked posts
No figure in this post describes measured A20 Pro silicon, because none exists yet. Apple's claims are labeled as Apple's claims, the allocation reporting is labeled as reporting, and the first-hand measurements are this site's own and are about a different chip.
