The paper argues that multimodal systems should be evaluated in settings that look more like real use. Multi-turn interaction can reveal failures that single image-question pairs miss, especially when models must update their understanding over time.
As vision-language models enter agents and assistants, these evaluations become more important. Real users rarely ask one perfect question; they iterate, correct and add context.