English
Related papers

Related papers: Enhancing In-the-Wild Speech Emotion Conversion wi…

200 papers

In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due to their deduction ability. In response to this problem, this…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Pengfei Wu , Junjie Pan , Chenchang Xu , Junhui Zhang , Lin Wu , Xiang Yin , Zejun Ma

Current approaches for controlling dialogue response generation are primarily focused on high-level attributes like style, sentiment, or topic. In this work, we focus on constrained long-term dialogue generation, which involves more…

Computation and Language · Computer Science 2022-05-17 Ramya Ramakrishnan , Hashan Buddhika Narangodage , Mauro Schilman , Kilian Q. Weinberger , Ryan McDonald

Accurately analyzing spontaneous, unconscious micro-expressions is crucial for revealing true human emotions, but this task remains challenging in wild scenarios, such as natural conversation. Existing research largely relies on datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yigui Feng , Qinglin Wang , Yang Liu , Ke Liu , Haotian Mo , Enhao Huang , Gencheng Liu , Mingzhe Liu , Jie Liu

For real-time speech enhancement (SE) including noise suppression, dereverberation and acoustic echo cancellation, the time-variance of the audio signals becomes a severe challenge. The causality and memory usage limit that only the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Chengyu Zheng , Yuan Zhou , Xiulian Peng , Yuan Zhang , Yan Lu

This work investigates how emotional speech and generative strategies affect ASR performance. We analyze speech synthesized from three emotional TTS models and find that substitution errors dominate, with emotional expressiveness varying…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Ya-Tse Wu , Chi-Chun Lee

Perceptual modification of voice is an elusive goal. While non-experts can modify an image or sentence perceptually with available tools, it is not clear how to similarly modify speech along perceptual axes. Voice conversion does make it…

Sound · Computer Science 2023-12-15 Robin Netzorg , Ajil Jalal , Luna McNulty , Gopala Krishna Anumanchipalli

Recently, Denoising Diffusion Probabilistic Models (DDPMs) have attained leading performances across a diverse range of generative tasks. However, in the field of speech synthesis, although DDPMs exhibit impressive performance, their long…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Xiangyu Zhang , Daijiao Liu , Hexin Liu , Qiquan Zhang , Hanyu Meng , Leibny Paola Garcia , Eng Siong Chng , Lina Yao

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the target language. Building such a system faces challenges of…

Sound · Computer Science 2023-10-09 Yuke Li , Xinfa Zhu , Yi Lei , Hai Li , Junhui Liu , Danming Xie , Lei Xie

Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Chanhyuk Choi , Taesoo Kim , Donggyu Lee , Siyeol Jung , Taehwan Kim

Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality severely limits the signal processing and understanding…

Computation and Language · Computer Science 2025-09-25 Jialong Mai , Xiaofen Xing , Yawei Li , Weidong Chen , Zhipeng Li , Jingyuan Xing , Xiangmin Xu

In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples containing both…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Vahid Noroozi , Zhehuai Chen , Somshubra Majumdar , Steve Huang , Jagadeesh Balam , Boris Ginsburg

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional TTS for seen…

Sound · Computer Science 2023-05-24 Minki Kang , Wooseok Han , Sung Ju Hwang , Eunho Yang

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While…

Sound · Computer Science 2026-02-13 Chung-Soo Ahn , Rajib Rana , Sunil Sivadas , Carlos Busso , Jagath C. Rajapakse

We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model…

Computation and Language · Computer Science 2024-06-28 Rendi Chevi , Alham Fikri Aji

Speech synthesis has recently seen significant improvements in fidelity, driven by the advent of neural vocoders and neural prosody generators. However, these systems lack intuitive user controls over prosody, making them unable to rectify…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-13 Max Morrison , Zeyu Jin , Justin Salamon , Nicholas J. Bryan , Gautham J. Mysore

We present EMPHASIS, an emotional phoneme-based acoustic model for speech synthesis system. EMPHASIS includes a phoneme duration prediction model and an acoustic parameter prediction model. It uses a CBHG-based regression network to model…

Audio and Speech Processing · Electrical Eng. & Systems 2018-06-27 Hao Li , Yongguo Kang , Zhenyu Wang

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

Social and Information Networks · Computer Science 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

Social interactions incorporate nonverbal signals to convey emotions alongside speech, including facial expressions and body gestures. Generative models have demonstrated promising results in creating full-body nonverbal animations…

Human-Computer Interaction · Computer Science 2026-04-01 Kiran Chhatre , Renan Guarese , Andrii Matviienko , Christopher Peters

Paraphrase generation, a.k.a. paraphrasing, is a common and important task in natural language processing. Emotional paraphrasing, which changes the emotion embodied in a piece of text while preserving its meaning, has many potential…

Computation and Language · Computer Science 2022-12-08 Justin Xie

Emotion recognition from speech is one of the key steps towards emotional intelligence in advanced human-machine interaction. Identifying emotions in human speech requires learning features that are robust and discriminative across diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-30 Alison Marczewski , Adriano Veloso , Nívio Ziviani