A new cs.AI paper, “What We are Missing in Multimodal LLM Evaluation?”, argues that evaluation has not kept pace with multimodal model capabilities.
The paper says many benchmarks test isolated skills, leaving open the harder question of whether a model can integrate evidence across modalities. That matters as systems move from captioning and visual QA toward richer assistants that combine text, image, audio, and video context.
The work highlights the need for evaluation designs that measure synthesis, not just task-specific accuracy.