English

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

Computer Vision and Pattern Recognition 2026-03-03 v3

Abstract

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, limiting research in video fusion and the inherent difficulty of jointly modeling spatial and temporal dependencies in a unified framework. To this end, we construct M3SVD, a benchmark dataset with 220220 temporally synchronized and spatially registered infrared-visible videos comprising 153,797153,797 frames, bridging the data gap. Secondly, we propose VideoFusion, a multi-modal video fusion model that exploits cross-modal complementarity and temporal dynamics to generate spatio-temporally coherent videos from multi-modal inputs. Specifically, 1) a differential reinforcement module is developed for cross-modal information interaction and enhancement, 2) a complete modality-guided fusion strategy is employed to adaptively integrate multi-modal features, and 3) a bi-temporal co-attention mechanism is devised to dynamically aggregate forward-backward temporal contexts to reinforce cross-frame feature representations. Experiments reveal that VideoFusion outperforms existing image-oriented fusion paradigms in sequences, effectively mitigating temporal inconsistency and interference. Project and M3SVD: https://github.com/Linfeng-Tang/VideoFusion.

Keywords

Cite

@article{arxiv.2503.23359,
  title  = {VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion},
  author = {Linfeng Tang and Yeda Wang and Meiqi Gong and Zizhuo Li and Yuxin Deng and Xunpeng Yi and Chunyu Li and Han Xu and Hao Zhang and Jiayi Ma},
  journal= {arXiv preprint arXiv:2503.23359},
  year   = {2026}
}

Comments

Accepted to CVPR 2026. The dataset and code are available at https://github.com/Linfeng-Tang/VideoFusion

R2 v1 2026-06-28T22:39:26.303Z