MAVIN: Multi-Shot Audio-Visual Generation with Narrative Control
Abstract
While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing methods suffer from temporal misalignment, limited controllability, and incomplete scripting. In this paper, we propose MAVIN, the first framework for multi-shot audio-visual generation with customized narrative control. To resolve temporal misalignment, we propose boundary-aware attention, which leverages hierarchical captions and boundary-aware token routing to render audio-visual elements within their respective temporal boundaries. To improve the controllability for multi-subject scenarios, we propose ID-aware propagation, utilizing identity embeddings and an identity-aware mask to bind specific identities to consistent visual appearances and vocal timbres. To provide comprehensive audio-visual narratives, we present a multi-agent scripting pipeline to transform free-form user inputs into hierarchical captions. Furthermore, we construct MAVINSet, a multi-shot audio-visual dataset for robust training and evaluation. Extensive experiments demonstrate that MAVIN achieves state-of-the-art performance, opening up a new avenue for integrating generative models into professional filmmaking workflows.
Cite
@article{arxiv.2606.29473,
title = {MAVIN: Multi-Shot Audio-Visual Generation with Narrative Control},
author = {Kaiqi Liu and Yunyao Mao and Ziqi Cai and Zheng Geng and Jing Wang and Qiulin Wang and Xintao Wang and Pengfei Wan and Kun Gai and Shuchen Weng and Boxin Shi},
journal= {arXiv preprint arXiv:2606.29473},
year = {2026}
}