Researchers introduce SportD, a benchmark for evaluating whether vision-language models can reason about physical strategy rather than only describe images. The tasks ask models to interpret dynamic scenes and make decisions that depend on action, timing, and spatial relationships.

That distinction matters as VLMs move into robotics, coaching, simulation, and embodied-agent settings. A system that can label what it sees may still fail when asked to infer what should happen next.

The benchmark adds another pressure test for multimodal models, emphasizing practical reasoning over static recognition scores.