English
Related papers

Related papers: emotion2vec: Self-Supervised Pre-Training for Spee…

200 papers

Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Hugo Bohy , Minh Tran , Kevin El Haddad , Thierry Dutoit , Mohammad Soleymani

Emotion detection in text is an important task in NLP and is essential in many applications. Most of the existing methods treat this task as a problem of single-label multi-class text classification. To predict multiple emotions for one…

Computation and Language · Computer Science 2019-11-11 Chenyang Huang , Amine Trabelsi , Xuebin Qin , Nawshad Farruque , Osmar R. Zaïane

We propose EmoDistill, a novel speech emotion recognition (SER) framework that leverages cross-modal knowledge distillation during training to learn strong linguistic and prosodic representations of emotion from speech. During inference,…

Computation and Language · Computer Science 2024-03-18 Debaditya Shome , Ali Etemad

This paper presents our contributions to the Speech Emotion Recognition in Naturalistic Conditions (SERNC) Challenge, where we address categorical emotion recognition and emotional attribute prediction. To handle the complexities of natural…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-15 Hyo Jin Jon , Longbin Jin , Hyuntaek Jung , Hyunseo Kim , Donghun Min , Eun Yi Kim

Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied. We present an embodied emotion classification dataset, CHEER-Ekman, extending the existing binary…

Computation and Language · Computer Science 2025-09-26 Phan Anh Duong , Cat Luong , Divyesh Bommana , Tianyu Jiang

Visual Emotion Analysis (VEA) aims to bridge the affective gap between visual content and human emotional responses. Despite its promise, progress in this field remains limited by the lack of open-source and interpretable datasets. Most…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yijie Guo , Dexiang Hong , Weidong Chen , Zihan She , Cheng Ye , Xiaojun Chang , Zhendong Mao

Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acoustic reconstruction, leaving emotion expressiveness…

Sound · Computer Science 2026-05-13 Jiacheng Shi , Hongfei Du , Xinyuan Song , Y. Alicia Hong , Yanfu Zhang , Ye Gao

Self-supervised pre-training paradigms have been extensively explored in the field of skeleton-based action recognition. In particular, methods based on masked prediction have pushed the performance of pre-training to a new height. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-03 Ruizhuo Xu , Linzhi Huang , Mei Wang , Jiani Hu , Weihong Deng

Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a framework for unifying SER tasks based on a multi-scale…

Sound · Computer Science 2022-07-29 Zuheng Kang , Junqing Peng , Jianzong Wang , Jing Xiao

Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech…

Sound · Computer Science 2025-11-12 Ziqian Zhang , Min Huang , Zhongzhe Xiao

Despite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous presence of all involved multiple modalities. Motivated by…

Machine Learning · Computer Science 2019-07-25 Jing Han , Zixing Zhang , Zhao Ren , Björn Schuller

Audio deepfake is so sophisticated that the lack of effective detection methods is fatal. While most detection systems primarily rely on low-level acoustic features or pretrained speech representations, they frequently neglect high-level…

Sound · Computer Science 2025-09-16 Xiaokang Li , Yicheng Gong , Dinghao Zou , Xin Cao , Sunbowen Lee

General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general tasks such as skip-gram models and natural language…

Computation and Language · Computer Science 2020-11-03 Aparna Khare , Srinivas Parthasarathy , Shiva Sundaram

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Tianhua Qi , Wenming Zheng , Björn W. Schuller , Zhaojie Luo , Haizhou Li

Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-03 Ismail Rasim Ulgen , Zongyang Du , Carlos Busso , Berrak Sisman

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts…

Machine Learning · Computer Science 2022-11-08 Itai Gat , Hagai Aronowitz , Weizhong Zhu , Edmilson Morais , Ron Hoory

Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Hao Cheng , Zhiwei Zhao , Yichao He , Zhenzhen Hu , Jia Li , Meng Wang , Richang Hong

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. Here, we examine how wav2vec2 models trained on four different…

Computation and Language · Computer Science 2025-06-13 Michele Gubian , Ioana Krehan , Oli Liu , James Kirby , Sharon Goldwater

Inferring emotion status from users' queries plays an important role to enhance the capacity in voice dialogues applications. Even though several related works obtained satisfactory results, the performance can still be further improved. In…

Sound · Computer Science 2018-10-26 Zefang Zong , Hao Li , Qi Wang