Recent AI models have made robots better at interpreting scenes and choosing actions, but researchers interviewed by MIT Technology Review say that progress is far from the general-purpose humanoids promoted by technology executives.

Vision-language-action models extend systems that understand images and text by producing physical commands. Google DeepMind’s Gemini Robotics, for example, can control a two-armed ALOHA 2 test platform to pack a lunchbox after training on demonstrations. That is a meaningful advance over writing separate rules for every movement.

The physical world remains much less forgiving than a chat interface. Robots must cope with changing objects, lighting, friction, balance and unexpected contact, while failures can damage property or hurt people. Training data is also harder and more expensive to collect than internet text or images. A machine shaped like a person is not necessarily able to learn the range of tasks a person performs.

Companies including Tesla and Nvidia have offered aggressive timelines for useful humanoids. Robotics specialists quoted in the report dispute that confidence and emphasize the gap between a controlled demonstration and sustained work in homes or factories. The technology is advancing, but current evidence does not support treating broad human-level dexterity as an imminent consumer product.