中文
相关论文

相关论文: emotion2vec: Self-Supervised Pre-Training for Spee…

200 篇论文

Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Hugo Bohy , Minh Tran , Kevin El Haddad , Thierry Dutoit , Mohammad Soleymani

Emotion detection in text is an important task in NLP and is essential in many applications. Most of the existing methods treat this task as a problem of single-label multi-class text classification. To predict multiple emotions for one…

计算与语言 · 计算机科学 2019-11-11 Chenyang Huang , Amine Trabelsi , Xuebin Qin , Nawshad Farruque , Osmar R. Zaïane

We propose EmoDistill, a novel speech emotion recognition (SER) framework that leverages cross-modal knowledge distillation during training to learn strong linguistic and prosodic representations of emotion from speech. During inference,…

计算与语言 · 计算机科学 2024-03-18 Debaditya Shome , Ali Etemad

This paper presents our contributions to the Speech Emotion Recognition in Naturalistic Conditions (SERNC) Challenge, where we address categorical emotion recognition and emotional attribute prediction. To handle the complexities of natural…

音频与语音处理 · 电气工程与系统科学 2025-10-15 Hyo Jin Jon , Longbin Jin , Hyuntaek Jung , Hyunseo Kim , Donghun Min , Eun Yi Kim

Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied. We present an embodied emotion classification dataset, CHEER-Ekman, extending the existing binary…

计算与语言 · 计算机科学 2025-09-26 Phan Anh Duong , Cat Luong , Divyesh Bommana , Tianyu Jiang

Visual Emotion Analysis (VEA) aims to bridge the affective gap between visual content and human emotional responses. Despite its promise, progress in this field remains limited by the lack of open-source and interpretable datasets. Most…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yijie Guo , Dexiang Hong , Weidong Chen , Zihan She , Cheng Ye , Xiaojun Chang , Zhendong Mao

Neural speech codecs provide discrete representations for speech language models, but emotional cues are often degraded during quantization. Existing codecs mainly optimize acoustic reconstruction, leaving emotion expressiveness…

声音 · 计算机科学 2026-05-13 Jiacheng Shi , Hongfei Du , Xinyuan Song , Y. Alicia Hong , Yanfu Zhang , Ye Gao

Self-supervised pre-training paradigms have been extensively explored in the field of skeleton-based action recognition. In particular, methods based on masked prediction have pushed the performance of pre-training to a new height. However,…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Ruizhuo Xu , Linzhi Huang , Mei Wang , Jiani Hu , Weihong Deng

Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a framework for unifying SER tasks based on a multi-scale…

声音 · 计算机科学 2022-07-29 Zuheng Kang , Junqing Peng , Jianzong Wang , Jing Xiao

Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech…

声音 · 计算机科学 2025-11-12 Ziqian Zhang , Min Huang , Zhongzhe Xiao

Despite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous presence of all involved multiple modalities. Motivated by…

机器学习 · 计算机科学 2019-07-25 Jing Han , Zixing Zhang , Zhao Ren , Björn Schuller

Audio deepfake is so sophisticated that the lack of effective detection methods is fatal. While most detection systems primarily rely on low-level acoustic features or pretrained speech representations, they frequently neglect high-level…

声音 · 计算机科学 2025-09-16 Xiaokang Li , Yicheng Gong , Dinghao Zou , Xin Cao , Sunbowen Lee

General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general tasks such as skip-gram models and natural language…

计算与语言 · 计算机科学 2020-11-03 Aparna Khare , Srinivas Parthasarathy , Shiva Sundaram

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Ziqi Zhang , Cheng Deng

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech…

音频与语音处理 · 电气工程与系统科学 2026-04-07 Tianhua Qi , Wenming Zheng , Björn W. Schuller , Zhaojie Luo , Haizhou Li

Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion…

音频与语音处理 · 电气工程与系统科学 2024-06-03 Ismail Rasim Ulgen , Zongyang Du , Carlos Busso , Berrak Sisman

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts…

机器学习 · 计算机科学 2022-11-08 Itai Gat , Hagai Aronowitz , Weizhong Zhu , Edmilson Morais , Ron Hoory

Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Hao Cheng , Zhiwei Zhao , Yichao He , Zhenzhen Hu , Jia Li , Meng Wang , Richang Hong

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. Here, we examine how wav2vec2 models trained on four different…

计算与语言 · 计算机科学 2025-06-13 Michele Gubian , Ioana Krehan , Oli Liu , James Kirby , Sharon Goldwater

Inferring emotion status from users' queries plays an important role to enhance the capacity in voice dialogues applications. Even though several related works obtained satisfactory results, the performance can still be further improved. In…

声音 · 计算机科学 2018-10-26 Zefang Zong , Hao Li , Qi Wang