A new arXiv paper proposes paper-grounded figure-to-video generation for scientific communication.

The task is to generate narrated walkthroughs that explain a complex figure step by step while grounding each part of the narration to regions of the image and the source paper.

The authors introduce MINARD and the FigTalk benchmark, pointing to a practical use case for multimodal generation: making dense scientific visuals easier to understand without losing connection to the underlying paper.