Google’s Gemini AI Now Controls Humanoid Robots for Complex Tasks
Google DeepMind released Gemini Robotics 2, an AI system that enables robots to perform dextrous tasks autonomously—from screwing lightbulbs to tidying…

Google DeepMind has released a new version of Gemini that does something previous AI models could not: control humanoid robots to perform intricate physical tasks without human intervention.
The system, called Gemini Robotics 2, combines multiple AI models into one framework. A vision language model processes images and video to understand the robot's surroundings and reason through tasks. Two vision language action models—trained specifically for physical movement—manage the robot's full-body motion and gripper control.
Video demonstrations showed robots performing work autonomously. Apptronik's Apollo 2 robot, fitted with hands from Sharpa, tidied shelves on its own. Another handled tasks like screwing in lightbulbs and tying trash bags. The model was trained using human teleoperation, video examples, and simulations.
"It's another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can," said Carolina Parada, head of robotics at Google DeepMind, in an interview with WIRED.
The capability marks a departure for Google. While Anthropic and OpenAI have dominated chatbots and coding tools, Google holds a stronger record in robotics research. The company previously partnered with Boston Dynamics, a leader in legged robots, to power those machines with AI brains. This release signals Google's broader bet that frontier AI models must move beyond software to unlock their full potential.
The training approach reveals current limitations. Complex tasks still require specific preparation—human demos, video examples, or simulated environments. AI models cannot yet perform a wide range of novel physical tasks without this targeted work. That constraint may shift as robotics datasets grow and training methods improve, but for now, each new capability demands engineering effort.
The move arrives as cloud providers race to embed AI into their platforms. Oracle announced Thursday it would add Gemini models to Oracle Cloud Infrastructure, joining offerings from Cohere, Meta, OpenAI, xAI, and Alibaba. The strategy reflects how enterprise customers now expect multiple AI options rather than a single proprietary model.
In automotive, General Motors is taking a different approach. GM plans to launch a proprietary in-vehicle AI assistant later this year that goes beyond the Gemini AI bot it deployed in millions of 2022-model-year vehicles and newer. "We'll be launching a more deeply integrated native AI assistant that combines conversational AI with GM vehicle knowledge and OnStar intelligence to create those capabilities that go beyond what a general purpose assistant can do," said Anna Santos, GM director of product management of voice and AI/machine learning, to CNBC.
The new GM assistant will integrate with vehicle telemetry and OnStar data to offer predictive maintenance and auto-focused features—tasks a general-purpose chatbot cannot handle. While Gemini lets drivers control temperature and radio settings through voice, the custom GM system aims deeper into the vehicle's operating knowledge.
These parallel advances show AI vendors pursuing different paths: Google pushing robotics as the next frontier, Oracle packaging multiple models for enterprise choice, and automakers building vertical solutions tailored to their hardware and data. Each strategy reflects the same conviction—AI's next growth lies not in conversation alone but in action.



