Google researchers have built an extended-reality prototype that lets conversational assistants point, demonstrate actions and emphasize warnings with virtual hands. AgentHands synchronizes those gestures with speech and anchors them to objects in a user’s physical surroundings.

Users first register nearby objects through eye gaze and scene reconstruction, producing a map of 3D bounding boxes. A language model then inserts gesture events into its response. Those events select motions such as pointing, tracing or miming a grip, while a local headset parser uses word-level timestamps to keep animation aligned with spoken instructions.

The system can also add visual cues. A hand may glow red to mark a hot surface, for example, or demonstrate where to pour. That makes it different from flat bounding boxes commonly placed over a camera feed: the guidance occupies the same three-dimensional space as the task.

In a study with 20 participants, the researchers report that AgentHands improved engagement, social presence and the ability to connect verbal directions with physical locations. It remains a research prototype presented at CHI 2026, rather than an announced consumer feature.