Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis
Abstract
Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.
Cite
@article{arxiv.2607.12686,
title = {Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis},
author = {Alexios Filippakopoulos and Elias Kallioras and Nikolaos Xiros and Efthymios Georgiou and Alexandros Potamianos},
journal= {arXiv preprint arXiv:2607.12686},
year = {2026}
}