MIT Technology Review: humanoid robot hype outruns what chatbot-style AI can deliver

MIT Technology Review: humanoid robot hype outruns what chatbot-style AI can deliver

MIT Technology Review, in collaboration with Aventine, a non-profit research foundation, argues that the loudest predictions about humanoid robots rest on a shaky assumption: that the AI advances behind ChatGPT and Claude will let robots imitate human movement the way chatbots imitate human language.

The feature opens with Tesla's Optimus, a white humanoid with a black head and torso that keeps turning up in videos. It dances, passes popcorn, vacuums and presses microwave buttons, but also falls backward while handing out water bottles and struggles to iron a shirt. Elon Musk, Tesla's CEO, believes Optimus will be "not just Tesla's biggest product ever, but probably the biggest product ever." He told shareholders in July that it "will have human and then superhuman dexterity" eventually, argues that it could automate almost all human labor for as little as $20,000 each, and predicted at Davos in January that it could be on sale to the public by the end of 2027. Marc Andreessen, cofounder and general partner of Andreessen Horowitz, has said robotics could become the "biggest industry in the history of the planet." In January, Nvidia CEO Jensen Huang said humanoid robots would match human-level ability this year. Morgan Stanley says the number of robots that "resemble and act like humans" is likely to reach nearly 1 billion by 2050, creating a market worth over $5 trillion.

Many robotics researchers push back. Yann LeCun, often called one of the godfathers of AI, said at Davos that none of the companies building humanoid robots has any idea how to make them smart enough to be useful. Jonathan Hurst, cofounder and chief robot officer of Agility Robotics and a robotics professor at Oregon State University, says looking like a person is easy; moving and behaving like one is dramatically harder. The article also says it is misleading to conflate humanoid robots with generalist machines that can learn and perform many tasks.

The piece then explains where real progress comes from, using Google DeepMind's work. Its researchers use ALOHA 2 ("A Low-cost Open-source Hardware System for Bimanual Teleoperation"), just a pair of arms, some grippers and a couple of cameras, to test their Gemini Robotics system. Controlled by it, the arms can pack a lunchbox: slide white bread into a Ziploc bag, close it, put grapes in a Tupperware container, secure the lid and zip everything into the box. The article calls this an objective step forward from what was possible three years ago.

The change is in robot policies, the part that decides how a robot reads its surroundings, plans movement and performs a task. These used to be thousands of lines of engineer-written rules. They are now handed to AI models, first vision-language models (VLMs), which are like large language models but trained on images as well as words, and then vision-language-action models (VLAs), which add motion commands. VLAs learn from images or video of a task plus data on how a robot arm moves, usually gathered by teleoperation. Gemini Robotics is a VLA trained on many hours of human demonstrations, and it can pick up snow peas with kitchen tongs, do origami or assemble a simple lunch. The limit: for now, a VLA-controlled robot asked to do something outside its training set is highly likely to fail. Edward Johns, a robotics professor at Imperial College London, says a real generalist policy would handle everything across the space of all tasks, while a Gemini Robotics model today can do only "a few things here and a few things there."

The usual answer is more data. Pannag Sanketi, a former tech lead in robotics at Google DeepMind who now works on his own AI robotics project, says DeepMind wants "as much data as possible" and favors a "multi-prong" approach. But there is no ocean of existing physical demonstrations, as there was text for language models, and each workaround has a flaw: teleoperation data is costly and time-consuming, training on videos of people yields poor-quality data, and deploying robots in the real world to collect data is unsafe and unreliable outside labs.

Hurst calls the belief that training data alone is the answer "a fundamentally flawed premise." Even making coffee involves endless variation: no two kitchens are identical, machines work differently, cups need different grips, and grounds, hot water and milk are each handled differently. Generality through VLAs would need "complete data coverage of all of the things that [a robot] could ever do," in effect an almost infinite pool of data. LeCun says approaches that worked for language do not work for high-dimensional, continuous, noisy data, and that "you have to use something else."

The article names world models as the leading contender for that something else. They are trained less on text than on video, three-dimensional scans and sensor data, and built to predict the outcomes of actions in the real world. The hope is faster, cheaper, safer development in simulations faithful to real physics, and robots that can anticipate the consequences of an action before taking it. Nvidia and Google are working on the technology, and investor money is flowing into startups. World Labs, cofounded by Stanford AI researcher Fei-Fei Li, raised $1 billion in February and was acquired by AMD at the end of September for $8.2 billion. AMI Labs, cofounded by LeCun, formerly Meta's chief AI scientist, also raised $1 billion in March. Even so, Li described the field late last year as "nascent," adding that "foundational approaches are still being established," and in a June Substack post she described daunting challenges. The available text breaks off while discussing the state of world models.

Key facts

  • Musk says Optimus could automate almost all human labor for as little as $20,000 each and predicted public sales by the end of 2027; Huang said in January that humanoids would match human-level ability this year.
  • Morgan Stanley estimates nearly 1 billion robots that resemble and act like humans by 2050, a market worth over $5 trillion.
  • Google DeepMind's Gemini Robotics, a vision-language-action model, can pack a lunchbox on ALOHA 2 hardware, but a VLA robot given a task outside its training set is highly likely to fail.
  • More training data is the standard answer, but teleoperation is costly, video of people gives poor data, and real-world deployment is unsafe outside labs; Jonathan Hurst calls data alone "a fundamentally flawed premise."
  • World models are the leading alternative: World Labs raised $1 billion in February and was acquired by AMD for $8.2 billion at the end of September, and LeCun's AMI Labs raised $1 billion in March.

Why it matters

Huge claims are attached to humanoid robots: a market worth over $5 trillion by 2050 per Morgan Stanley, human-level ability this year per Huang, public sales of Optimus by the end of 2027 per Musk. The article's point is that these forecasts lean on one idea, that the methods behind chatbots will carry over to the physical world. Roboticists quoted in it, including LeCun and Hurst, doubt that, and the piece frames the open question as whether current AI methods are enough or an entirely new path is required. It also separates two things the hype blurs: robots that look like people and generalist machines that can learn many tasks.

Who it affects

Investors and companies betting on humanoid robots, including Tesla, Nvidia and Google, are the most directly affected, along with the startups raising money for world models. Robotics researchers face the choice between scaling data for VLAs and pursuing world models. For ordinary readers, the headline's message is that the robot revolution is not imminent, even though the article says progress in labs is real if painstaking.

How to use it

This is analysis, not a product or tool. It gives readers a way to judge robot demos and forecasts: ask whether a task sits inside the robot's training data, since a VLA-controlled robot is highly likely to fail outside it. A tidy lunchbox demo and a viral video of a humanoid falling over are both consistent with the article's picture of real but limited capability. It also gives a vocabulary worth knowing: robot policies, VLMs, VLAs and world models.

How solid is it

It is a feature in MIT Technology Review, produced with Aventine, and it rests on named sources: Hurst, Johns, Sanketi and LeCun, plus public statements from Musk, Huang and Andreessen. The forecasts and funding figures are stated as reported claims, and the article presents no new experiment or result of its own. The portion of the text available ends partway through the discussion of world models, so the conclusion is not covered here.

Risks and caveats

The skeptics give no timeline for when generalist humanoids might become useful, and the article says roboticists disagree on exactly when a robotics revolution could arrive. The proponents' figures are predictions: Musk's price and dates, Huang's claim and Morgan Stanley's estimate are all forward-looking. The data-versus-new-approach debate is unresolved: Sanketi favors a multi-prong data strategy, while Hurst and LeCun reject data alone. World models are described by Li herself as nascent, with daunting challenges, so they are a leading contender rather than a proven fix.

“It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.”

— Jonathan Hurst, cofounder and chief robot officer of Agility Robotics and professor of robotics at Oregon State University