English
Related papers

Related papers: ED-TTS: Multi-Scale Emotion Modeling using Cross-D…

200 papers

Although automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-15 Wen Wu , Chao Zhang , Philip C. Woodland

Cross-lingual Speech Emotion Recognition (CLSER) aims to identify emotional states in unseen languages. However, existing methods heavily rely on the semantic synchrony of complete labels and static feature stability, hindering low-resource…

Sound · Computer Science 2026-04-10 Ya Zhao , Yinfeng Yu , Liejun Wang

Pre-trained models (PTMs) have shown great promise in the speech and audio domain. Embeddings leveraged from these models serve as inputs for learning algorithms with applications in various downstream tasks. One such crucial task is Speech…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-25 Orchid Chetia Phukan , Arun Balaji Buduru , Rajesh Sharma

As an essential element for the diagnosis and rehabilitation of psychiatric disorders, the electroencephalogram (EEG) based emotion recognition has achieved significant progress due to its high precision and reliability. However, one…

Machine Learning · Computer Science 2021-07-19 Hao Chen , Ming Jin , Zhunan Li , Cunhang Fan , Jinpeng Li , Huiguang He

Automatic speech emotion recognition (SER) by a computer is a critical component for more natural human-machine interaction. As in human-human interaction, the capability to perceive emotion correctly is essential to take further steps in a…

Sound · Computer Science 2022-10-27 Bagus Tris Atmaja , Masato Akagi

This paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). As a pre-trained model, it can learn prosody attributes from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-15 Liping Chen , Yan Deng , Xi Wang , Frank K. Soong , Lei He

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text…

Sound · Computer Science 2026-02-12 Raymond Chung

Cross-lingual speech emotion recognition (SER) is important for a wide range of everyday applications. While recent SER research relies heavily on large pretrained models for emotion training, existing studies often concentrate solely on…

Sound · Computer Science 2024-07-09 Shreya G. Upadhyay , Carlos Busso , Chi-Chun Lee

Massively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results…

Computation and Language · Computer Science 2024-05-31 Zhuoyuan Mao , Chenhui Chu , Sadao Kurohashi

Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yanchao Zhao , Jihao Zhu , Yu Liu , Weizhuo Chen , Yuling Yang , Kun Peng

Emotion is a complicated notion present in music that is hard to capture even with fine-tuned feature engineering. In this paper, we investigate the utility of state-of-the-art pre-trained deep audio embedding methods to be used in the…

Sound · Computer Science 2021-04-15 Eunjeong Koh , Shlomo Dubnov

Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in…

Sound · Computer Science 2025-06-02 Yi Chang , Zhao Ren , Zhonghao Zhao , Thanh Tam Nguyen , Kun Qian , Tanja Schultz , Björn W. Schuller

Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Rui Liu , Pu Gao , Jiatian Xi , Berrak Sisman , Carlos Busso , Haizhou Li

With the rapid advancement in deep generative models, recent neural Text-To-Speech(TTS) models have succeeded in synthesizing human-like speech. There have been some efforts to generate speech with various prosody beyond monotonous prosody…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-24 Seongho Joo , Hyukhun Koh , Kyomin Jung

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

Social and Information Networks · Computer Science 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotion-Aware Stepwise…

Computation and Language · Computer Science 2026-02-10 Jiacheng Shi , Hongfei Du , Yangfan He , Y. Alicia Hong , Ye Gao

In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts. On the one hand, an LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-04 Jaejin Cho , Raghavendra Pappagari , Purva Kulkarni , Jesus Villalba , Yishay Carmiel , Najim Dehak

The performance of most emotion recognition systems degrades in real-life situations ('in the wild' scenarios) where the audio is contaminated by reverberation. Our study explores new methods to alleviate the performance degradation of SER…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Ohad Cohen , Gershon Hazan , Sharon Gannot

Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a framework for unifying SER tasks based on a multi-scale…

Sound · Computer Science 2022-07-29 Zuheng Kang , Junqing Peng , Jianzong Wang , Jing Xiao