English

M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment

Computer Vision and Pattern Recognition 2024-03-15 v1 Multimedia Sound Audio and Speech Processing

Abstract

This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling.

Keywords

Cite

@article{arxiv.2403.09451,
  title  = {M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment},
  author = {Long Nguyen-Phuoc and Renald Gaboriau and Dimitri Delacroix and Laurent Navarro},
  journal= {arXiv preprint arXiv:2403.09451},
  year   = {2024}
}
R2 v1 2026-06-28T15:20:12.689Z