English

InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection

Multimedia 2025-09-25 v1 Sound

Abstract

Detecting emotional inconsistency across modalities is a key challenge in affective computing, as speech and text often convey conflicting cues. Existing approaches generally rely on incomplete emotion representations and employ unconditional fusion, which weakens performance when modalities are inconsistent. Moreover, little prior work explicitly addresses inconsistency detection itself. We propose InconVAD, a two-stage framework grounded in the Valence/Arousal/Dominance (VAD) space. In the first stage, independent uncertainty-aware models yield robust unimodal predictions. In the second stage, a classifier identifies cross-modal inconsistency and selectively integrates consistent signals. Extensive experiments show that InconVAD surpasses existing methods in both multimodal emotion inconsistency detection and modeling, offering a more reliable and interpretable solution for emotion analysis.

Keywords

Cite

@article{arxiv.2509.20140,
  title  = {InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection},
  author = {Zongyi Li and Junchuan Zhao and Francis Bu Sung Lee and Andrew Zi Han Yee},
  journal= {arXiv preprint arXiv:2509.20140},
  year   = {2025}
}

Comments

5 pages, 1 figure, 3 tables

R2 v1 2026-07-01T05:54:11.547Z