One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids
Summary
Keerthana reads China’s viral Robot Olympics as genuine locomotion progress but a poor proxy for near-term robot utility. Running on flat, rigid terrain is unusually simulation-friendly, while useful manipulation breaks down around contact, friction, deformable objects, and unpredictable environments. Her blunt calibration: “Many of my friends are very productive, but we don’t run faster than Usain Bolt.”
Robotics remains at its “GPT-2 moment” because impressive demonstrations have not yet become portable, reliable intelligence. GPT-3-level progress would require learning from multiple examples across many tasks and a brain that transfers across humanoids, arms, and other bodies; today, competence still depends heavily on the exact robot and setup. One-shot imitation is promising, but copying a demonstration is not the same as generalizing when the scene changes.
Gemini Robotics 2 separates embodied reasoning from physical execution across three models. Gemini Robotics ER2 is the higher-level, tool-using “brain”; Gemini Robotics 2 is the vision-language-action model controlling motion; and Gemini Robotics On-Device 2 is a smaller local version. The architecture lets a large model interpret intent while specialized action models handle bodies, but additional model handoffs create latency and failure risks.
Controlled commercial environments remain the likeliest early deployment wedge, even though model latency may matter less at home. Nathan emphasizes repeatable factory and warehouse tasks and professional servicing; Keerthana agrees that commercial use cases are more controlled than homes. A six-second pause is untenable on a conveyor belt but irrelevant if a robot folds laundry overnight. Keerthana expects fast and slow models to develop together, with capabilities transferred from larger systems into smaller ones.
Cross-embodiment is improving rapidly, but “one brain, any body” is still an aspiration rather than a solved capability. Google DeepMind’s models now control humanoids “from fingertips to feet,” and its on-device program reportedly adapts tasks to new bodies with roughly 200 examples. Yet Keerthana sees little precedent for taking a new embodiment fully zero-shot to many tasks at very high reliability.
Hardware has moved faster than Keerthana expected, shifting the remaining work toward reliability, repeatability, cost, and safe force control, while dexterous hands approach a capability bottleneck. In roughly a year and a quarter, the frontier advanced from grippers to multifinger control, garbage-bag tying, and whole-body manipulation. The Shadow hand, she says, may lift about 20 kg. The remaining commercial work is making hands repeatable, durable, affordable, and safe around delicate objects and people.
Keerthana refuses to declare either world models or any single data source the winning robotics recipe. Teleoperation is accurate but expensive and difficult to scale; instrumented UMI demonstrations scale better while retaining precise sensor-derived action labels; egocentric human video is abundant but noisy. Her recurring principle is empirical: forecasts should be tested against results rather than treated as settled, because robotics remains early enough that several architectures and data mixtures may prove useful.
Deep dive
1. Viral robot races showcase the easiest part of embodiment
Keerthana first encountered China’s Robot Olympics through her father and friends in India, who were struck by humanoids apparently running faster than the fastest people. The demonstrations captured the public imagination because running is an intuitive benchmark—and because robots have progressed dramatically since DARPA-era machines struggled to walk through doors without falling.
Nathan’s pushback—worth keeping—is whether speed is useful rather than merely spectacular. Keerthana allows that fast locomotion might matter in defense, but her team is focused on robots helping people in the physical world: “Many of my friends are very productive, but we don’t run faster than Usain Bolt.” Leg speed is rarely the limiting factor in valuable work.
Locomotion has advanced faster than manipulation because flat ground and rigid bodies can be modeled relatively well. Controllers can learn motion in simulation, often through imitation learning followed by reinforcement-learning fine-tuning, and transfer it to a physical robot; contact-rich tasks introduce harder physics. Picking up a wooden cube is tractable, while folding cloth exposes friction, deformation, and interactions that modern simulators still represent poorly.
2. The research frontier has moved from grasping to whole-body intelligence
Keerthana remembers when basic pick-and-place was a serious research problem. Foundation models made many of those behaviors comparatively easy, pushing researchers toward humanoids, multifinger hands, and coordinated whole-body manipulation. Her team began working on humanoids around two years earlier, when even grasping an apple was difficult.
Gemini Robotics 2’s meaningful advance is not simply walking or moving quickly; it can squat, reposition itself, and manipulate unfamiliar objects in a closed loop. Keerthana points to a video in which a humanoid changes posture while trying to retrieve a watering can—behavior she would have found surprising just a year earlier.
“Whole-body control” means the vision-language-action model considers the humanoid from fingertips to feet while processing the surrounding scene. Lower-level mechanisms may still contribute stabilization, but the semantic controller is not merely issuing commands to an isolated arm: it coordinates bodily actions in service of a user’s task.
3. Reliability and generalization remain separate axes
Keerthana frames skill and generalization as “orthogonal axes.” A narrowly engineered system can already perform one or two tasks extremely well, but the cost of executing many narrow tasks can become roughly the same as the cost of executing the first or second. A general foundation raises performance across many tasks and makes subsequent specialization cheaper, echoing the transition from task-specific language models to GPT- and Gemini-style foundations.
Nathan presses for a METR-like curve connecting task complexity to 50%, 80%, or 99% success. Keerthana does not offer a clean curve, but says pick-and-place is approaching a strong, usable signal. Her larger point is that robotics cannot optimize only for task success: it must also preserve the broad understanding that supports adaptation.
Physical deployment sharply raises the reliability bar. A language model can omit something and be corrected; a robot can break an object, create a hazard, or damage itself. Tasks with cheap retries are therefore much easier: a robot can undo a LEGO mistake, but after dropping an egg it has created a slippery mess and a second problem to solve.
4. One-shot demonstrations are prompts, not proof of generality
Recent systems can learn from a handful of demonstrations—or apparently one—but Keerthana places these results on a continuum with ordinary prompting. A foundation model may be guided by language, an image with an indicated object, or a video demonstration. In-context learning reduces deployment time because it avoids task-specific retraining, but it does not abolish the underlying generalization problem.
The critical test is what happens after the exact demonstration changes. These models can memorize and reproduce what they were shown; the harder question is whether they adapt to a different scene, object, or arrangement. Pick-and-place benefits from abundant prior data, while “show it once, then tie a garbage bag” remains a much stronger challenge.
Keerthana therefore still calls robotics “GPT-2.” Reaching a GPT-3 analogue would require learning from multiple examples across many tasks and substantially stronger transfer across bodies. A brain that works only on one robot in one configuration is not yet a general brain in the way software behaves consistently across phones, computers, and operating systems.
5. Gemini Robotics 2 divides reasoning, action, and local execution
Keerthana describes Gemini Robotics ER2 as a tool-using embodied-reasoning system built on the Gemini Flash models, oriented toward understanding images, video, semantics, and robotic situations. Gemini Robotics 2 is the VLA that converts user intent or ER2 guidance into action, while Gemini Robotics On-Device 2 is a smaller model designed to run locally on the robot.
Nathan says his coding agent connected ER2 to a browser-based 2D scene with simple grasp-and-place tools. The model understood requests such as moving a banana to a plate. He also tested an object that moved before grasping; the model failed, and then received another image with the object in a different place. In his small test, Flash and ER2 performed mostly the same overall, and both did well.
ER2 was released via API first because it is closer to the mature Gemini stack. Action models remain farther from broad deployment and require more research into reliability and serving. Opening ER2 beyond verified partners also produced a much larger stream of feedback, including academic benchmarks and experiments using it for unexpectedly low-level control.
6. Latency and memory force a hierarchy of robotic brains
Nathan observed about six seconds between asking ER2 to move an object and receiving a tool call. Keerthana’s answer is workload-specific: a conveyor cannot wait, but a homeowner may not care how slowly laundry is folded overnight. Larger models may unlock advanced capabilities while smaller or on-device models supply responsiveness, local execution, and operation without dependable cloud connectivity.
She expects fast and slow systems to coexist rather than arrive sequentially. Some abilities may emerge first in large models and then be transferred into smaller ones. Cloud deployments can run on distant servers—for example, in Michigan—while on-device models address applications that need local execution.
Nathan describes ER2’s 128,000-token context as roughly three minutes of densely represented robotic history, depending on frame sampling and tokenization. Keerthana’s context-engineering principle is “very dense memory”: retain raw visual detail when it matters, but summarize long activity into compact text when the exact frame is less important. A cooking robot may need to remember that it flipped an egg, not preserve every pixel of the motion.
7. Rich interfaces matter more than a universal robot API today
ER2 can be given tools much like a digital Gemini agent, but robotics lacks a standard interface equivalent to the Model Context Protocol. Every body exposes different controls and APIs, so unfamiliar embodiments may require additional engineering. Keerthana says current robot-specific models remain better at direct control for now, with larger reasoning models orchestrating them.
The interface between ER2 and a VLA is already broader than raw coordinates. A VLA can respond to language, pointing, images, or video demonstrations; ER2 can instruct it to look left, lower its head, return, or interact with a referenced object. Keerthana’s direction of travel is a richer connection in which text is only one communication modality.
Nathan’s banana-versus-lemon example exposes the design question: should ER2 name the object, point to it, describe its shape, or provide pixel coordinates? Keerthana does not declare one permanent answer because “controllability” is an active research area. The appropriate abstraction will evolve with what action models can reliably understand.
8. Cross-embodiment gains are real, but orchestration compounds errors
Google DeepMind is developing VLAs with Apptronik, Agile Robots, and Boston Dynamics. Keerthana says the field has made “a big jump” in cross-embodiment transfer, yet fully zero-shot deployment on a new body across many tasks at high reliability still lacks a meaningful precedent.
The on-device program demonstrates the value of a common physical foundation: with roughly 200 examples, testers can teach multiple tasks on different bodies because the model already understands much of the physical world and has prior control experience. Most embodiments nevertheless cluster around familiar designs—hands similar to human hands, vectors in space, bimanual hands, and humanoids—while hands mounted on drones sit in the exotic long tail.
Nathan points out that sequencing ten subtasks would create compounding opportunities for error. Keerthana emphasizes that completion detection, transitions, and deciding when to replan can fail even when reasoning and motor execution work well in isolation.
9. Hands are improving faster than their economics
In roughly a year and a quarter, Keerthana updated her own hardware forecast. Tasks once near the frontier for grippers can increasingly be performed by multifinger hands, while those hands add abilities grippers never had—garbage-bag tying, precise finger motion, and manipulation integrated with posture.
The market now offers several capable hands, but reliability, repeatability, and price remain unresolved. The Wuji hand, Keerthana says, is comparable to a ten-year-old child’s: human-sized but somewhat weak. The Shadow hand may lift about 20 kg and has been seen opening jars, though torque and success on tightly sealed lids remain uncertain.
Soft robotics matters where compliance and tactile feedback improve manipulation, but Keerthana again centers practical criteria: “What is repeatable, and what is durable?” Instrumented gloves and UMI-style systems are interesting because they can capture demonstrations and sensor-derived action information; force sensing also lets models join delicate parts without crushing them.
10. Safety is a capability spanning mechanics, perception, and intent
Keerthana rejects a simple capability-versus-safety tradeoff: “People will not use dangerous robots and dangerous agents.” The most useful systems must follow instructions without causing chaos or violating constraints. Nathan and Keerthana both expect commercial sites to precede homes; Keerthana emphasizes that commercial environments are more controlled, while homes contain children and carry a higher safety bar.
Robotics adds operational failures distinct from malicious behavior. A humanoid can injure someone simply by falling, losing a sensor, or misjudging contact. In one test, someone put a basket over the robot’s head; the right response was not to continue blindly, but to recognize the obstruction and request its removal.
Safety mechanisms range from hard emergency stops that cut power—potentially making an unstable robot collapse—to soft stops that freeze it while maintaining balance. Better force sensing and compliant control add another layer. For Keerthana, safety is a “full-system thing,” extending from mechanical design and emergency controls to perception, planning, and high-level adherence to human intent.
11. Humanoid interaction, world models, and data recipes remain open
Gemini Robotics 2 introduced more natural human-robot interaction: gestures are generated in context rather than individually scripted. During filming, a robot could discuss its own task, acknowledge where it might have erred, and gesture while conversing. Keerthana calls this “semi-emergent”—deliberately developed as an interaction capability, but not programmed gesture by gesture.
Human likeness raises both attachment and expectations. New observers can forget that a humanoid is a machine, while judging its mistakes more harshly than those of an obviously mechanical arm. Keerthana argues humanoids are scientifically valuable because they expose three frontiers simultaneously: whole-body control, multifinger dexterity, and natural human-robot interaction.
On predictive world models, her honest answer is that “the jury is still out.” Robotics researchers should be emotionally attached to the problem, not a favorite method; VLAs, action models, and explicit forecasting systems may all contribute. The field’s “recipes are not stabilized,” which makes empirical comparison more valuable than architectural certainty.
The same agnosticism governs data. Teleoperation is precise but costly and difficult to scale, and may become less useful as robots develop; sensor-rich UMI demonstrations are more scalable while retaining action labels; egocentric human video is abundant but noisy about exact hand and end-effector motion. Keerthana expects a mixture that changes as hardware improves, with users ultimately needing both high-level delegation and low-level control when personalization or safety demands it.