Close Menu
    National News Brief
    Tuesday, September 15
    • Home
    • Business
    • Lifestyle
    • Science
    • Technology
    • International
    • Arts & Entertainment
    • Sports
    National News Brief
    Home » Google DeepMind wants Gemini to power many different robots

    Google DeepMind wants Gemini to power many different robots

    Team_NationalNewsBriefBy Team_NationalNewsBriefSeptember 15, 2026 Science No Comments12 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email


    When Kanishka Rao was a kid, robots like the droids from Star Wars and Rosie from The Jetsons got pretty far into his head. And now that he’s a principal software engineer at Google DeepMind, Rao is trying to get pretty far into theirs.

    Rao remembers those bots as “helpful around the house but also sassy.” Today, in an office surrounded by galumphing, fidgeting robots of all kinds, Rao and DeepMind are trying to at least make them helpful. Any chatbot can simulate sass; Rao’s aiming to build general-purpose intelligence that can inhabit many different robot bodies—in pursuit of what roboticists call “physical AI.”

    In late July, DeepMind showed off its latest attempt, Gemini Robotics 2. The previous generation controlled person-shaped robots mostly from the waist up, but the new software can control whole-body movement—legs, torso, arms, fingers. It works across machines ranging from two-armed research platforms to Apptronik’s Apollo 2 humanoid, letting robots fetch snacks on command, change lightbulbs, tie knots. They’re not our robot overlords yet, but they’re starting to act like robot servants.


    On supporting science journalism

    If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


    The lightbulb is almost beside the point. DeepMind is trying to bring to robotics one of the tricks that make large language models so powerful: feed a model enough data that what it learns in one place still works somewhere else. Here that means carrying a skill from one job, or even one kind of robot body, to the next.

    Google is hardly alone in chasing this possibility. The capital and enthusiasm sloshing through artificial intelligence have spilled over into robotics. A bullish 2025 Morgan Stanley report projected a humanoid market worth upward of $5 trillion by 2050; Elon Musk has made Tesla’s Optimus humanoid a primary focus of the company. BMW has deployed humanoid robots in one of its factories. China, meanwhile, is packed with robot start-ups; one called Unitree is launching a $900-million IPO, having run several of its newest models through U.S. certification about a month before the Federal Communications Commission barred new foreign-built humanoids and robot dogs from the country. An extraordinary amount of money, and now a fenced-off market, is riding on machines that still struggle with everyday chores.

    A lot of that money hinges on “dexterous manipulation.” The robots you’ve seen doing backflips or kung fu have mastered their own gross mobility. Folding laundry or making scrambled eggs means mastering everything they touch. “The key distinction between the backflips and the eggs is this: One of them requires you to deeply understand yourself, your own body,” Rao says. “The other one requires you to understand the world.” That, he says, “is why manipulation is so hard. It’s about how you interact with the world.”


    A decade ago people building robots didn’t talk about “training” the way they do now. Robots were machines capable of complicated sets of motions, useful mainly in highly constrained environments such as automotive assembly lines or specific parts of warehouses. The algorithms that controlled the actions of robotic arms or bodies worked fine—until someone tried to put a robot into any of the messy, unstructured, unpredictable spaces we humans have built for our bodies and ourselves. Then the robot became just another dangerous piece of machinery. “The physical space sometimes needs to be precise to, like, centimeter, millimeter precision because they’re not intelligent,” says Carolina Parada, vice president and head of robotics at Google DeepMind. “All they’re doing is repeating motions.”

    Machine learning offered a way around all that painstaking instruction. Instead of specifying every motion, roboticists could effectively turn the machine loose and let it work out what to do on its own. DeepMind had already made its reputation this way with AlphaGo, the system it built to master the fiendishly complex board game Go. AlphaGo initially learned from thousands of games played by humans, then improved by playing versions of itself repeatedly, using reinforcement learning. Google acquired DeepMind in 2014; the next year AlphaGo beat European champion Fan Hui, and a year after that it defeated Lee Sedol, one of the game’s great players. In 2023 Google combined the division with the Google Brain Team to create Google DeepMind, the organization that would build Gemini.

    Today reinforcement learning is one way to get a robot to acquire new skills: Let it try to do the thing, whatever it is, in the environment where it’d have to do it (or a digital simulation), over and over. Whenever it does the right thing, it gets a little numerical reward. “It’s a massively powerful paradigm because you no longer have to show it how to do the task,” says Matei Ciocarlie, a roboticist at Columbia University. “But it takes a very long time.”

    There’s a quicker way. Humans can “demonstrate” the right movements, often by teleoperating the robot or using GoPro-like cameras to record themselves performing the task—the gig worker’s–eye view of the job. Simulation can add still more examples. This “imitation learning” approach gives data-acquiring robots a leg up (if they have legs), although it also has the disadvantage of humiliating us human meat bags while we train our replacements.

    Tying a trash bag is a deceptively hard test of the dexterous manipulation robot hands need for everyday chores.

    But imitation learning is predicated on the robots’ digital brains being able to fit the data into the context of their own bodies and capabilities. The robot has to be able to emulate a human’s demonstration with its own peculiar collection of joints. Sounds hard, but engineers working on this project have an analogy: language. Train a big enough model on enough varied data, and it can pick up patterns that transfer to things it was never explicitly taught.

    “We have an existence proof that a general-purpose model can control different robot morphologies because that’s what humans do. If you drive a car, once you’re proficient it feels like an extension of yourself,” says James Marshall, director of the Center for Machine Intelligence at the University of Sheffield in England. (He’s also co-founder of Opteran, a company that’s taking a whole other path—reverse engineering insect brains.) “So it’s not surprising that there could be a general technology that could control different morphologies. But moving from a quadruped to a humanoid or to a drone is more challenging and data-intensive because we don’t have a full understanding of how the brain solves that problem yet.”

    Gemini was already multimodal—a model family built with a knack for extracting useful information from words and images. And, of course, Gemini has already seen an absurd amount of the Internet. “There’s a lot of information Gemini already has about how the world works,” Parada says. But “vanilla Gemini models don’t know what it feels like to translate it into action. What we’re doing is teaching them.”

    So when one Googler sends a toddler-size bot to fetch a bag of popcorn from a nearby lounge—Google is famous for good snacks—think of what’s actually going on behind that bot’s eyes, in an entire stack of software. A higher-level reasoning model, Gemini Robotics ER 2, can break down a verbal request to get popcorn into a series of steps; a lower-level model turns what the robot sees and is told into the movements needed to carry the steps out, updating predictions of what’s going to happen next four or five times every second. ER 2 also watches the task unfold and estimates its progress, classifying each moment into one of five ranges of completion. That sounds almost comically basic until you consider how often a robot can execute a perfectly reasonable motion and still fail its overall objective.

    Eventually the robot does bring the popcorn back to the researcher. He has to kind of pry it from the robot’s cold, unliving pincer, but it mostly works. DeepMind reported 57.4 percent accuracy on the progress-classification test, which is better than the results from the models it tested against but nowhere near omniscience.

    The DeepMind researchers know they’re not quite there yet. “Our bet is that if you give [the robot] enough data, intelligence will emerge. That’s the same for language and for motion,” Rao says. “With enough data of the robot interacting with things around it, it will build this implicit thing such that it’ll be able to deal with generalizing to dexterous tasks.”

    They’ve also focused on safety, with protections intended to keep bots from accidentally hurting nearby humans—by having ER 2 bring a robot to a safe stop if someone gets too close, for example. The company also created a new benchmark called Asimov, after the science-fiction writer who famously created three laws to govern robot behavior, which suggests DeepMind is worried about on-purpose hurting, too. Gemini, though, provides one layer of safety; the robot bodies generally come with protections of their own. Texas robotics company Apptronik installs all kinds of sensors and safety systems into its Apollo to stop it from doing anything catastrophically stupid.

    In one Google video, a humanoid robot with a stylized face haltingly brushes detritus from a countertop into a dustpan. In another, a robot puts grapes into a plastic bag. In DeepMind’s own tests, Apollo pulled off the dustpan task just 32 percent of the time. In engineering terms, the bag and the brush bristles are “deformable”; the grapes are simply fragile. The robot has to manipulate objects that fold, bend and bruise.

    An extraordinary amount of money—and now a fenced-off market—rides on machines that still struggle with everyday chores.

    That’s especially important because this type of system, a so-called vision-language-action model, relies on visual input. These robots see a lot—they have more cameras than you or I have eyes—but as far as the Gemini model is concerned, they feel literally nothing. A robotic hand can have 22 degrees of freedom—that’s 22 independent ways to move—and still have almost no sense of what it’s touching. When a Gemini-powered robot lifts those grapes, it gets no fingertip sensation telling it when one is beginning to burst. When it picks up a wine glass, it can’t “feel” how much pressure the glass can take before it shatters or perceive the minimum amount of pressure to exert so that the glass won’t slip through its fingers. Even if we humans aren’t conscious of it, we have tactility and motor coordination wired into not only our brains but also our distal neurons and muscles, tendons, joints and digits. The digital digits don’t have our evolutionary advantages. “Humans are able to get feedback much sooner and be much more reactive,” Parada says. “That is part of what we’re constantly trying to improve.”

    Cutting-edge robot hardware often uses strain gauges and torque sensors to supply exactly that kind of feedback, but then there’s another issue. “The Internet has massive amounts of visual data—gigantic amounts—and it has essentially no tactile data or force data or proprioceptive data,” says Ciocarlie, who also co-founded the robot-hand company Tangent Robotics. Any given real-world task has both semantic and somatic elements. The tasks DeepMind has its models and robots working on are in some ways more intricate than plenty of jobs robots already do reliably. But the tasks DeepMind bots can do still don’t require human levels of dexterity, Ciocarlie says—“the kind of things where the semantic, conscious intelligence needs to be supplemented by motor intelligence.”

    Much of the new money in robotics is chasing machines that are person-shaped. This goal makes some sense; a machine that moves around on caterpillar treads and has 20 multi-degree-of-freedom tentacles bursting from an eight-foot-tall torso might be better at getting places and carrying stuff, but it would be perhaps less capable of operating in an environment built for humans, where, for example, countertops are usually about 25 inches deep and doorways are about 36 inches wide. Ironically, imitation learning has a reinforcement effect here, too. If a human is demonstrating or teleoperating a task, the robot is more likely to learn it if its own parts aren’t too different from the human’s—you want that “embodiment gap” to be as small as possible. It’s as if future robots will inherit technical debt from evolution itself.

    The term “embodiment” carries philosophical baggage, and here the irony doubles back. For decades proponents of the idea of embodied cognition argued against disembodied intelligence. Understanding, in this view, depends on having a body that moves through and responds to the physical world. (Another approach to robot control called a world model—touted by chipmaker Nvidia, among other laboratories—leans heavily on this idea.) By this logic, knowing what an apple is takes a lot more than a calculation of how the word “apple” relates to other words in a multidimensional vector space. It’s the fruit’s feel, its smell and taste, and a person’s preference for Galas over Fujis.

    It might be true. But Rao, at least, isn’t convinced. “Before I joined robotics, I was on the speech-recognition team, and I used to work on language modeling. I thought this would be true where surely you can’t know what an apple is until you’ve held an apple and tasted it,” he says. “I think I was totally wrong. It’s been the other way around. It’s the digital AIs that have really made the physical AI more powerful. It seems like you don’t need to touch all these objects or interact with them or see what they weigh to understand them.”

    About five years ago D. E. Wittkower, a philosopher at Old Dominion University in Virginia, borrowed philosopher Thomas Nagel’s 1974 question about the inner life of a bat for an essay called “What Is It Like to Be a Bot?” DeepMind’s robots suggest almost the opposite question—or at least ask a different one: Who cares? Maybe the machine does not need anything remotely like our experience of an apple to know enough about apples to handle one. Maybe these robots don’t need an inner life. But they could use some nerve endings.



    Source link

    Team_NationalNewsBrief
    • Website

    Keep Reading

    Math Puzzle: A puzzling picture

    The science of friendship and loneliness

    25 years after 9/11, AI sparks a math debate as data center pollution concerns grow

    Trump repeals emissions regulations of fossil-fuel power plants

    25 winners of math’s ‘Nobel Prize’ decry the AI invasion of their discipline

    What caused this woman’s hallucinations? Doctors reveal an infestation of brain worms

    Add A Comment

    Comments are closed.

    Editors Picks

    Bill Maher Calls Out Far Left Actor Sean Penn for Saying He Wouldn’t Meet Trump After Meeting Castro and Hugo Chavez (VIDEO) | The Gateway Pundit

    June 17, 2025

    Can it still matter if it doesn’t scale?

    October 31, 2025

    OG ‘Real Housewives Of New York’ Set For New Reality Show On E!

    February 4, 2026

    Nurses’ warnings on Trump’s tax and policy bill were too long ignored

    July 23, 2025

    Venezuela’s political transition talks wrap first day in Caracas

    August 7, 2026
    Categories
    • Arts & Entertainment
    • Business
    • International
    • Latest News
    • Lifestyle
    • Opinions
    • Politics
    • Science
    • Sports
    • Technology
    • Top Stories
    • Trending News
    • World Economy
    About us

    Welcome to National News Brief, your one-stop destination for staying informed on the latest developments from around the globe. Our mission is to provide readers with up-to-the-minute coverage across a wide range of topics, ensuring you never miss out on the stories that matter most.

    At National News Brief, we cover World News, delivering accurate and insightful reports on global events and issues shaping the future. Our Tech News section keeps you informed about cutting-edge technologies, trends in AI, and innovations transforming industries. Stay ahead of the curve with updates on the World Economy, including financial markets, economic policies, and international trade.

    Editors Picks

    Inside the Inference Hardware Revolution Of 2026

    September 15, 2026

    2026 World Economic Conference Tickets!

    September 15, 2026

    Eastern Mass. Division 5 Football Teams Gear Up for 2026 Season

    September 15, 2026

    Fans Can’t Get Over Catherine O’Hara’s Emmy Honor

    September 15, 2026
    Categories
    • Arts & Entertainment
    • Business
    • International
    • Latest News
    • Lifestyle
    • Opinions
    • Politics
    • Science
    • Sports
    • Technology
    • Top Stories
    • Trending News
    • World Economy
    • Privacy Policy
    • Disclaimer
    • Terms and Conditions
    • About us
    • Contact us
    Copyright © 2024 Nationalnewsbrief.com All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.