中文
相关论文

相关论文: Versatile audio-visual learning for emotion recogn…

200 篇论文

Effective human-agent interaction (HAI) relies on accurate and adaptive perception of human emotional states. While multimodal deep learning models - leveraging facial expressions, speech, and textual cues - offer high accuracy in emotion…

机器学习 · 计算机科学 2025-12-15 Matvey Nepomnyaschiy , Oleg Pereziabov , Anvar Tliamov , Stanislav Mikhailov , Ilya Afanasyev

Multimodal large models have been recognized for their advantages in various performance and downstream tasks. The development of these models is crucial towards achieving general artificial intelligence in the future. In this paper, we…

声音 · 计算机科学 2023-09-12 Sen Fang , Bowen Gao , Yangjian Wu , Teik Toe Teoh

Accurate emotion understanding in videos necessitates effectively recognizing and interpreting emotional states by integrating visual, textual, auditory, and contextual cues. Although recent Large Multimodal Models (LMMs) have exhibited…

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we extend this paradigm by leveraging motion and audio that…

计算机视觉与模式识别 · 计算机科学 2022-07-18 Rui Qian , Yeqing Li , Zheng Xu , Ming-Hsuan Yang , Serge Belongie , Yin Cui

Understanding human emotions from multimodal signals poses a significant challenge in affective computing and human-robot interaction. While multimodal large language models (MLLMs) have excelled in general vision-language tasks, their…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Xiaojiang Peng , Jingyi Chen , Zebang Cheng , Bao Peng , Fengyi Wu , Yifei Dong , Shuyuan Tu , Qiyu Hu , Huiting Huang , Yuxiang Lin , Jun-Yan He , Kai Wang , Zheng Lian , Zhi-Qi Cheng

Multimodal Emotion Recognition refers to the classification of input video sequences into emotion labels based on multiple input modalities (usually video, audio and text). In recent years, Deep Neural networks have shown remarkable…

机器学习 · 计算机科学 2024-10-28 Ashish Ramayee Asokan , Nidarshan Kumar , Anirudh Venkata Ragam , Shylaja S Sharath

In recent years, deep learning has achieved innovative advancements in various fields, including the analysis of human emotions and behaviors. Initiatives such as the Affective Behavior Analysis in-the-wild (ABAW) competition have been…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Seongjae Min , Junseok Yang , Sangjun Lim , Junyong Lee , Sangwon Lee , Sejoon Lim

Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained Vision-Language models like CLIP have shown remarkable semantic alignment capabilities, their…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Byeongjin Jung , Chanyeong Park , Sejoon Lim

Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational…

计算机视觉与模式识别 · 计算机科学 2021-05-13 Rameswar Panda , Chun-Fu Chen , Quanfu Fan , Ximeng Sun , Kate Saenko , Aude Oliva , Rogerio Feris

While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underexplored. We introduce WAVE (\textbf{u}nified \&…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Changli Tang , Qinfan Xiao , Ke Mei , Tianyi Wang , Fengyun Rao , Chao Zhang

Recently, researchers have gradually realized that in some cases, the self-supervised pre-training on large-scale Internet data is better than that of high-quality/manually labeled data sets, and multimodal/large models are better than…

声音 · 计算机科学 2023-08-08 Sen Fang , Yangjian Wu , Bowen Gao , Jingwen Cai , Teik Toe Teoh

Dimensional representations of speech emotions such as the arousal-valence (AV) representation provide a continuous and fine-grained description and control than their categorical counterparts. They have wide applications in tasks such as…

音频与语音处理 · 电气工程与系统科学 2024-02-07 Enting Zhou , You Zhang , Zhiyao Duan

We used two multimodal models for continuous valence-arousal recognition using visual, audio, and linguistic information. The first model is the same as we used in ABAW2 and ABAW3, which employs the leader-follower attention. The second…

多媒体 · 计算机科学 2023-04-18 Su Zhang , Ziyuan Zhao , Cuntai Guan

In the field of affective computing, traditional methods for generating emotions predominantly rely on deep learning techniques and large-scale emotion datasets. However, deep learning techniques are often complex and difficult to…

人机交互 · 计算机科学 2025-03-24 Haidong Wang , Qia Shan , JianHua Zhang , PengFei Xiao , Ao Liu

Multimodal emotion recognition has attracted much attention recently. Fusing multiple modalities effectively with limited labeled data is a challenging task. Considering the success of pre-trained model and fine-grained nature of emotion…

计算与语言 · 计算机科学 2023-03-02 Junyi He , Meimei Wu , Meng Li , Xiaobo Zhu , Feng Ye

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…

Automatic emotion recognition (ER) has recently gained lot of interest due to its potential in many real-world applications. In this context, multimodal approaches have been shown to improve performance (over unimodal approaches) by…

计算机视觉与模式识别 · 计算机科学 2022-09-20 R Gnana Praveen , Eric Granger , Patrick Cardinal

Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Jun Yu , Naixiang Zheng , Guoyuan Wang , Yunxiang Zhang , Lingsi Zhu , Jiaen Liang , Wei Huang , Shengping Liu

In the domain of human-computer interaction, accurately recognizing and interpreting human emotions is crucial yet challenging due to the complexity and subtlety of emotional expressions. This study explores the potential for detecting a…

多媒体 · 计算机科学 2025-05-13 Jiehui Jia , Huan Zhang , Jinhua Liang

Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate…

音频与语音处理 · 电气工程与系统科学 2023-05-22 Qiushi Zhu , Long Zhou , Ziqiang Zhang , Shujie Liu , Binxing Jiao , Jie Zhang , Lirong Dai , Daxin Jiang , Jinyu Li , Furu Wei