Robotics has spent most of the last decade being the AI subfield that was always five years away. That framing stopped being accurate sometime in 2025. Two things happened in the same twelve-month window that don't usually happen together in robotics: real foundation models shipped — open, benchmarked, buildable-on — and a company deployed enough of the resulting hardware, at production scale, to report hard numbers instead of pilot-program anecdotes. Here's what's actually documented, what's genuinely still unsolved, and where I'm declining to repeat a number I couldn't independently verify.
The foundation models: GR00T N1 and Gemini Robotics
On March 18, 2025, at GTC, Jensen Huang unveiled Nvidia Isaac GR00T N1 — the first open, downloadable humanoid robot foundation model, meaning any team can pull the weights and fine-tune on their own hardware rather than starting from nothing. Its architecture is a dual-system design borrowed conceptually from how researchers describe human cognition: a vision-language-model "System 2" handles reasoning about what to do, while a faster diffusion-transformer "System 1" handles the actual motor control. Huang's framing at the announcement — "the age of generalist robotics is here" — was aspirational, but the model backing it up was a real, inspectable release, not a teaser.
Six days earlier, on March 12, 2025, Google DeepMind had shipped its own answer: Gemini Robotics and a companion model, Gemini Robotics-ER, both built on top of Gemini 2.0's reasoning capability. An On-Device variant followed on June 24, 2025, notable for being fine-tunable on as few as 50 to 100 demonstrations — a meaningful drop in the data a team needs to specialize the model for a new task. Google DeepMind's launch partners for hardware included Apptronik (maker of the Apollo humanoid), Agility Robotics (Digit), and Boston Dynamics, giving the software stack immediate physical bodies to run on rather than shipping into a vacuum.
Solving the data problem with simulation
Robotics has always had a harder data problem than language models — you can't scrape a warehouse's worth of physical trial-and-error off the internet the way you can scrape text. Nvidia's answer is a simulation pipeline built around Cosmos, a family of "world foundation models" (including the 7-billion-parameter Cosmos Transfer) launched across CES and GTC 2025, paired with a physics engine called Newton, co-developed with Google DeepMind and Disney Research and also announced at GTC 2025.
The concrete output of that pipeline is the number I find most striking in this entire story: Nvidia's GR00T N1 blueprint generated 780,000 synthetic robot trajectories — the rough equivalent of 6,500 hours of human demonstration data, compressed from what would normally take about nine months down to eleven hours of generation time. Nvidia reports that training on this synthetic data on top of real demonstrations lifted model performance by 40% compared to training on real data alone. That 40% figure is Nvidia's own stated result from its own methodology, worth attributing plainly rather than presenting as an independently audited benchmark — but the underlying mechanism (simulation compressing months of data collection into hours) is the actual structural unlock, regardless of the exact performance delta.
Simulation's Data Leverage
Nvidia — GR00T N1 blueprint, arXiv 2503.14734
9 months → 11 hrs
Time to generate 780,000 synthetic trajectories (≈6,500 human-demo hours)
+40%
Performance lift vs. real-data-only training (Nvidia's own reported result)
Nvidia — GR00T N1 blueprint, arXiv 2503.14734
GR00T N1's Training Data Mix
Illustrative of the described data hierarchy, not precise measured proportions
Illustrative of the described data hierarchy, not precise measured proportions — Nvidia, arXiv 2503.14734
Proof of deployment: Amazon's one-millionth robot
Foundation models and simulation pipelines are upstream infrastructure; the number that actually proves something shipped is Amazon's. In July 2025, Amazon delivered its one-millionth warehouse robot — the specific unit went to a facility in Japan — coordinated by DeepFleet, a generative foundation model Amazon built to manage fleet logistics, which improved overall robot travel efficiency by 10%. Amazon now runs robots across more than 300 facilities, and by its own reporting, 75% of Amazon deliveries now involve a robot at some stage of the process. In May 2025, Amazon also unveiled Vulcan, a robot with touch and force-feedback sensing — a meaningfully harder capability than the pick-and-place robots that dominated the previous generation, since it lets a robot handle irregularly shaped or fragile items without pre-programmed handling rules.
This is the part of the story that separates 2025 from every prior "year of robotics" prediction: it isn't a pilot program or a single flagship facility. It's a fleet, at a scale where a 10% efficiency improvement is worth reporting as a headline metric on its own.
The 2025 Foundation-Model Timeline
Nvidia; Google DeepMind; Amazon
- Mar 12 2025Gemini Robotics launchesbuilt on Gemini 2.0 (Google DeepMind)
- Mar 18 2025GR00T N1 unveiledfirst open humanoid foundation model (Nvidia, GTC)
- Jun 24 2025Gemini Robotics On-Devicefine-tunable on 50-100 demonstrations
- Jul 2025Amazon: 1M robots + DeepFleet+10% fleet travel efficiency
Capital is pricing this as real
Money is a decent proxy for how seriously an industry believes its own hype, and the robotics capital markets moved fast in 2025. Figure AI raised a Series C of more than $1 billion at a $39 billion post-money valuation, announced September 16, 2025 and led by Parkway VC. That's roughly a 15x jump from its Series B — $675 million at a $2.6 billion valuation, in February 2024 — in about eighteen months. Figure has raised roughly $1.9 billion in total and has stated plans to build 100,000 humanoid robots over four years. Whether that production target is achievable is genuinely an open question; the valuation trajectory at least tells you institutional investors are underwriting the bet at scale, not hedging it.
I'll flag one thing I'm deliberately not including: specific benchmark-accuracy percentages for vision-language-action models on standard test suites. Nvidia's own paper describes GR00T N1 as outperforming prior imitation-learning baselines, which I'm comfortable repeating qualitatively — but I couldn't independently verify a specific accuracy number I'd be comfortable putting in print, so I'm leaving it as a qualitative claim rather than inventing precision the sourcing doesn't support. Same goes for Tesla's Optimus program — it's a real competitor in this space worth naming, but I'm not citing specific Gen 3 hardware specs here, because I couldn't verify them to the standard the rest of this piece is held to.
Figure AI's Valuation Trajectory
PR Newswire (Figure AI); TechCrunch
~15x jump in ~18 months, on ~$1.9B total raised — PR Newswire (Figure AI); TechCrunch
What's still unsolved
The honest caveat, and it's one the field states about itself rather than one critics are imposing from outside, is the sim-to-real gap: a policy trained almost entirely in simulation still has to survive contact with an open-ended real world full of states no simulator fully anticipated. Reliability at the scale Amazon has achieved is still the exception, not the industry norm — most humanoid and mobile-manipulation deployments outside a handful of companies remain pilot-scale. It's also worth being precise about Nvidia's own justification for the technology: the company frames embodied AI against a global labor shortage it puts at more than 50 million people worldwide. That's Nvidia's stated framing for why this matters economically, not an independently audited labor statistic, and it's worth reading it that way — as the company's argument for the market it's building toward, not as settled fact. The compute all of this runs on, incidentally, draws from the same inference-economics story reshaping AI infrastructure more broadly — a warehouse robot fleet is, among other things, a fleet of edge inference endpoints.
The view from Kathmandu
Labor migration defines Nepal's economy in a way it simply doesn't for most of the countries writing the headlines about embodied AI. A technology that manufactures physical labor at scale isn't an abstract Silicon Valley story from here — it's directly adjacent to the remittance economy that a huge share of Nepali households depend on. I don't think warehouse robots are coming for that specific labor market anytime soon; Amazon's fleet is inside Amazon's own logistics network, not a general-purpose labor replacement deployed globally. But the direction of travel — foundation models plus simulation plus falling hardware costs, compounding year over year — is the kind of curve worth watching closely from an economy built on exporting the exact kind of physical labor this technology is explicitly trying to automate. It connects to a broader pattern I keep circling back to in multi-agent coordination systems more generally: the interesting shift isn't any single model, it's what happens once fleets of agents — digital or physical — start coordinating with each other by default.
Sources
- Nvidia — Isaac GR00T N1 announcement, GTC, March 18, 2025 (nvidianews.nvidia.com)
- arXiv 2503.14734 — GR00T N1 technical paper, architecture and synthetic-data blueprint results
- Google DeepMind — Gemini Robotics and Gemini Robotics-ER launch, March 12, 2025; On-Device variant, June 24, 2025 (deepmind.google)
- arXiv 2503.20020 — Gemini Robotics technical report
- Amazon — "About Amazon," one-millionth warehouse robot and DeepFleet announcement, July 2025; Vulcan unveiling, May 2025
- TechCrunch; CNBC — coverage of Amazon's robot-fleet milestone, July 2025
- PR Newswire (Figure AI) — Series C, $39B valuation, September 16, 2025; TechCrunch — Figure Series B, February 2024
- Hugging Face Blog — Nvidia Cosmos world foundation models overview, CES/GTC 2025
Written by Abhishek Kushwaha, founder and writer at Global Tech Search, based in Kathmandu, Nepal.
