Google’s Gemini Now Controls Humanoid Robots—And Costs Matter
Google DeepMind released Gemini Robotics 2, enabling humanoids to perform complex physical tasks autonomously. Meanwhile, DeepSeek's cheaper AI model is…

Google DeepMind's latest Gemini model can now control humanoid robots to perform intricate physical tasks—from screwing in lightbulbs to tying trash bags. But the same week, a rival AI model from China undercut Google's pricing by orders of magnitude, signaling a widening competitive rift in enterprise artificial intelligence.
Gemini Robotics 2 combines multiple AI systems into one framework. A vision language model processes images and video to understand environments and reason through tasks. Two separate vision language action models—trained specifically for physical movement—control a robot's full-body motion and gripper precision. Google DeepMind demonstrated the system controlling Apptronik's Apollo 2 robot and other humanoids to complete real-world work.
The company trained these models using human teleoperation, video examples, and simulations. It's a pragmatic limitation: AI models cannot yet perform a wide range of complex tasks without specialized training data.
"It's another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can," said Carolina Parada, head of robotics at Google DeepMind.
The robotics breakthrough underscores Google's long-term bet that frontier AI must move beyond digital interfaces to unlock value. It previously partnered with Boston Dynamics, the legged-robot specialist, to provide AI "brains" for those machines.
The Price Shock
On the same day—August 3—a cheaper alternative emerged. DeepSeek's V4-Flash model, backed by China's Alibaba Group, cost just $0.03 to complete each benchmark test while scoring 50 on Artificial Analysis' Intelligence Index. That matches Google's Gemini 3.6 Flash's performance level.
The catch: those three cents represent benchmark costs, not typical customer queries. DeepSeek's actual pricing to users is $0.14 per million input tokens and $0.28 per million output tokens. Still cheap enough to pressure enterprise margins.
More capable models from OpenAI, Anthropic, and Moonshot scored at least seven points higher on the index—but their costs remain undisclosed or significantly steeper. The benchmark combines nine tests across coding, reasoning, and workplace tasks.
Alibaba's cloud division has already integrated DeepSeek V4 into its AI Gateway, allowing customers to route workloads between DeepSeek and Alibaba's own Qwen models. On August 3, Alibaba also unveiled Qwen3.8-Max, a 2.4-trillion-parameter model, sending its Hong Kong shares up 7%.
Google's Stronger Position
For Alphabet, the threat is real but containable. Google Cloud revenue jumped 82 percent to $24.8 billion in Q2 2026, with operating margin expanding to 35.6 percent. That margin growth signals customers still value Google's integrated infrastructure, distribution, and broader product ecosystem—not just the AI model itself.
Alibaba faces harder arithmetic. Its cloud revenue grew 38 percent in the March quarter, and AI products now represent 30 percent of external cloud sales. But group revenue rose only 3 percent overall. Alibaba is also committing to exceed its three-year RMB380 billion AI investment pledge while treating margins as secondary.
Lower inference prices can pull new customers into Alibaba Cloud. Revenue compounds only if workload volume grows faster than prices fall—and those customers also purchase storage, networking, and databases alongside compute.
For enterprises, the timing is clear: AI model costs are collapsing. What happens next depends on whether suppliers can expand into adjacent services faster than pricing erodes margins.



