English
Related papers

Related papers: Investigating Polyglot Speech Foundation Models fo…

200 papers

Data-intensive fine-tuning of speech foundation models (SFMs) to scarce and diverse dysarthric and elderly speech leads to data bias and poor generalization to unseen speakers. This paper proposes novel structured speaker-deficiency…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Shujie Hu , Xurong Xie , Mengzhe Geng , Jiajun Deng , Zengrui Jin , Tianzi Wang , Mingyu Cui , Guinan Li , Zhaoqing Li , Helen Meng , Xunying Liu

Emotion recognition in conversations (ERC) is a rapidly evolving task within the natural language processing community, which aims to detect the emotions expressed by speakers during a conversation. Recently, a growing number of ERC methods…

Computation and Language · Computer Science 2023-12-12 Tao Shi , Xiao Liang , Yaoyuan Liang , Xinyi Tong , Shao-Lun Huang

Large language models (LLMs) showcase increasingly impressive English benchmark scores, however their performance profiles remain inconsistent across multilingual settings. To address this gap, we introduce PolyPrompt, a novel,…

Computation and Language · Computer Science 2025-06-04 Nathan Roll

This paper addresses the formulation of a new speaker identification approach which employs knowledge of emotional content of speaker information. Our proposed approach in this work is based on a two-stage recognizer that combines and…

Sound · Computer Science 2018-01-23 Ismail Shahin

Large pre-trained models are essential in paralinguistic systems, demonstrating effectiveness in tasks like emotion recognition and stuttering detection. In this paper, we employ large pre-trained models for the ACM Multimedia Computational…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-17 Dejan Porjazovski , Yaroslav Getman , Tamás Grósz , Mikko Kurimo

Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making…

Sound · Computer Science 2026-03-11 Hezhao Zhang , Huang-Cheng Chou , Shrikanth Narayanan , Thomas Hain

Speech Emotion Recognition is a crucial area of research in human-computer interaction. While significant work has been done in this field, many state-of-the-art networks struggle to accurately recognize emotions in speech when the data is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Rashedul Hasan , Meher Nigar , Nursadul Mamun , Sayan Paul

Usually, people talk neutrally in environments where there are no abnormal talking conditions such as stress and emotion. Other emotional conditions that might affect people talking tone like happiness, anger, and sadness. Such emotions are…

Sound · Computer Science 2017-07-04 Ismail Shahin

In the Emotion Recognition in Conversation task, recent investigations have utilized attention mechanisms exploring relationships among utterances from intra- and inter-speakers for modeling emotional interaction between them. However,…

Computation and Language · Computer Science 2024-09-24 Jieying Xue , Minh Phuong Nguyen , Blake Matheny , Le Minh Nguyen

Emotion recognition plays a vital role in enhancing human-computer interaction. In this study, we tackle the MER-SEMI challenge of the MER2025 competition by proposing a novel multimodal emotion recognition framework. To address the issue…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Juewen Hu , Yexin Li , Jiulin Li , Shuo Chen , Pring Wong

Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches…

Sound · Computer Science 2024-06-14 Xueyu Liu , Jie Lin , Chao Wang

This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-23 Liang-Yeh Shen , Shi-Xin Fang , Yi-Cheng Lin , Huang-Cheng Chou , Hung-yi Lee

Cross-lingual speech emotion recognition (SER) is important for a wide range of everyday applications. While recent SER research relies heavily on large pretrained models for emotion training, existing studies often concentrate solely on…

Sound · Computer Science 2024-07-09 Shreya G. Upadhyay , Carlos Busso , Chi-Chun Lee

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable multi- and cross-modal integration capabilities. However, their potential for fine-grained emotion understanding remains systematically underexplored.…

Human-Computer Interaction · Computer Science 2025-12-25 Jing Han , Zhiqiang Gao , Shihao Gao , Jialing Liu , Hongyu Chen , Zixing Zhang , Björn W. Schuller

Despite recent advancements in speech emotion recognition (SER) models, state-of-the-art deep learning (DL) approaches face the challenge of the limited availability of annotated data. Large language models (LLMs) have revolutionised our…

Sound · Computer Science 2024-06-21 Siddique Latif , Muhammad Usama , Mohammad Ibrahim Malik , Björn W. Schuller

Speech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep…

Sound · Computer Science 2023-05-12 Yi Chang , Zhao Ren , Thanh Tam Nguyen , Kun Qian , Björn W. Schuller

Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF…

Sound · Computer Science 2025-08-07 Bartłomiej Marek , Piotr Kawa , Piotr Syga

While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presents UASR-LLM, a novel framework that adapts frozen SFMs to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jing-Xuan Zhang , Genshun Wan , Jin Li , Jianqing Gao , Duo Zhao , Zhen-Hua Ling

Recent developments in speech emotion recognition (SER) often leverage deep neural networks (DNNs). Comparing and benchmarking different DNN models can often be tedious due to the use of different datasets and evaluation protocols. To…

Sound · Computer Science 2021-10-08 Neil Scheidwasser-Clow , Mikolaj Kegler , Pierre Beckmann , Milos Cernak

Speech separation aims to separate multiple speech sources from a speech mixture. Although speech separation is well-solved on some existing English speech separation benchmarks, it is worthy of more investigation on the generalizability of…

Sound · Computer Science 2022-03-14 Kuan-Po Huang , Yuan-Kuei Wu , Hung-yi Lee