English
Related papers

Related papers: Empathy Level Prediction in Multi-Modal Scenario w…

200 papers

This paper studies compressing pre-trained language models, like BERT (Devlin et al.,2019), via teacher-student knowledge distillation. Previous works usually force the student model to strictly mimic the smoothed labels predicted by the…

Computation and Language · Computer Science 2020-05-11 Xing Wu , Yibing Liu , Xiangyang Zhou , Dianhai Yu

Human beings have rich ways of emotional expressions, including facial action, voice, and natural languages. Due to the diversity and complexity of different individuals, the emotions expressed by various modalities may be semantically…

Artificial Intelligence · Computer Science 2023-02-06 Chuan Zhang , Daoxin Zhang , Ruixiu Zhang , Jiawei Li , Jianke Zhu

Learning based on multimodal data has attracted increasing interest recently. While a variety of sensory modalities can be collected for training, not all of them are always available in development scenarios, which raises the challenge to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Shicai Wei , Yang Luo , Chunbo Luo

Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great promise by learning deep representations aligned across…

In this study, we focus on automated approaches to detect depression from clinical interviews using multi-modal machine learning (ML). Our approach differentiates from other successful ML methods such as context-aware analysis through…

Machine Learning · Computer Science 2024-12-30 Genevieve Lam , Huang Dongyan , Weisi Lin

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder…

Artificial Intelligence · Computer Science 2023-12-15 Haoyu Zhang , Yu Wang , Guanghao Yin , Kejun Liu , Yuanyuan Liu , Tianshu Yu

Audio-Video Emotion Recognition is now attacked with Deep Neural Network modeling tools. In published papers, as a rule, the authors show only cases of the superiority in multi-modality over audio-only or video-only modality. However, there…

Signal Processing · Electrical Eng. & Systems 2021-08-02 Xin Chang , Władysław Skarbek

This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Distillation (KD). The widely adopted mutual information…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Mengxi Chen , Linyu Xing , Yu Wang , Ya Zhang

How are emotions formed? Through extensive debate and the promulgation of diverse theories , the theory of constructed emotion has become prevalent in recent research on emotions. According to this theory, an emotion concept refers to a…

Artificial Intelligence · Computer Science 2024-10-27 Kazuki Tsurumaki , Chie Hieida , Kazuki Miyazawa

Multi-modal affect recognition models leverage complementary information in different modalities to outperform their uni-modal counterparts. However, due to the unavailability of modality-specific sensors or data, multi-modal models may not…

Image and Video Processing · Electrical Eng. & Systems 2021-08-03 Vandana Rajan , Alessio Brutti , Andrea Cavallaro

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wangyuan Zhu , Jun Yu

The integration of different imaging modalities, such as structural, diffusion tensor, and functional magnetic resonance imaging, with deep learning models has yielded promising outcomes in discerning phenotypic characteristics and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-08 Zhiyuan Li , Hailong Li , Anca L. Ralescu , Jonathan R. Dillman , Mekibib Altaye , Kim M. Cecil , Nehal A. Parikh , Lili He

Learning Using Privileged Information is a particular type of knowledge distillation where the teacher model benefits from an additional data representation during training, called privileged information, improving the student model, which…

Computation and Language · Computer Science 2024-08-20 Rafael-Edy Menadil , Mariana-Iuliana Georgescu , Radu Tudor Ionescu

Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of integrating temporal and/or spatial attention mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2017-03-13 Chiori Hori , Takaaki Hori , Teng-Yok Lee , Kazuhiro Sumi , John R. Hershey , Tim K. Marks

Prompt learning has emerged as a valuable technique in enhancing vision-language models (VLMs) such as CLIP for downstream tasks in specific domains. Existing work mainly focuses on designing various learning forms of prompts, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Zheng Li , Xiang Li , Xinyi Fu , Xin Zhang , Weiqiang Wang , Shuo Chen , Jian Yang

Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most existing research assume that all modalities are available during both training and testing,…

Sound · Computer Science 2026-04-21 Weide Liu , Huijing Zhan

Modeling text-based time-series to make prediction about a future event or outcome is an important task with a wide range of applications. The standard approach is to train and test the model using the same input window, but this approach…

Computation and Language · Computer Science 2023-01-27 Jinghui Liu , Daniel Capurro , Anthony Nguyen , Karin Verspoor

We propose a cross-modal attention distillation framework to train a dual-encoder model for vision-language understanding tasks, such as visual reasoning and visual question answering. Dual-encoder models have a faster inference speed than…

Computation and Language · Computer Science 2022-10-18 Zekun Wang , Wenhui Wang , Haichao Zhu , Ming Liu , Bing Qin , Furu Wei

Supervisory signals have the potential to make low-dimensional data representations, like those learned by mixture and topic models, more interpretable and useful. We propose a framework for training latent variable models that explicitly…

Multimodal emotion recognition in conversations aims to infer utterance-level emotions by jointly modeling textual, acoustic, and visual cues within context. Despite recent progress, key challenges remain, including redundant cross-modal…

Sound · Computer Science 2026-04-17 Chengling Guo , Yuntao Shou , Tao Meng , Wei Ai , Yun Tan , Keqin Li
‹ Prev 1 4 5 6 7 8 10 Next ›