English
Related papers

Related papers: Can Emotion Fool Anti-spoofing?

200 papers

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on…

Sound · Computer Science 2025-05-27 Zhichao Wu , Yueteng Kang , Songjun Cao , Long Ma , Qiulin Li , Qun Yang

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the…

Computation and Language · Computer Science 2022-07-14 Yookyung Shin , Younggun Lee , Suhee Jo , Yeongtae Hwang , Taesu Kim

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not…

Computation and Language · Computer Science 2023-12-20 Rui Liu , Yifan Hu , Yi Ren , Xiang Yin , Haizhou Li

Self-supervised speech model is a rapid progressing research topic, and many pre-trained models have been released and used in various down stream tasks. For speech anti-spoofing, most countermeasures (CMs) use signal processing algorithms…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-07 Xin Wang , Junichi Yamagishi

The major aim of this paper is to explain the data poisoning attacks using label-flipping during the training stage of the electroencephalogram (EEG) signal-based human emotion evaluation systems deploying Machine Learning models from the…

It is now well-known that automatic speaker verification (ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofing countermeasure…

Cryptography and Security · Computer Science 2024-01-30 Xuechen Liu , Md Sahidullah , Kong Aik Lee , Tomi Kinnunen

Large Language Models like GPT-4 adjust their responses not only based on the question asked, but also on how it is emotionally phrased. We systematically vary the emotional tone of 156 prompts - spanning controversial and everyday topics -…

Computation and Language · Computer Science 2025-07-30 Franck Bardol

Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing…

Sound · Computer Science 2025-09-03 Arnab Das , Yassine El Kheir , Carlos Franzreb , Tim Herzig , Tim Polzehl , Sebastian Möller

Emotional voice conversion models adapt the emotion in speech without changing the speaker identity or linguistic content. They are less data hungry than text-to-speech models and allow to generate large amounts of emotional data for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-15 Bastian Schnell , Goeric Huybrechts , Bartek Perz , Thomas Drugman , Jaime Lorenzo-Trueba

Speech emotion recognition (SER) systems often exhibit gender bias. However, the effectiveness and robustness of existing debiasing methods in such multi-label scenarios remain underexplored. To address this gap, we present EMO-Debias, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-06 Yi-Cheng Lin , Huang-Cheng Chou , Yu-Hsuan Li Liang , Hung-yi Lee

The data scarcity problem in Electroencephalography (EEG) based affective computing results into difficulty in building an effective model with high accuracy and stability using machine learning algorithms especially deep learning models.…

Machine Learning · Computer Science 2021-09-09 Zhi Zhang , Sheng-hua Zhong , Yan Liu

Empathetic response generation endeavors to empower dialogue systems to perceive speakers' emotions and generate empathetic responses accordingly. Psychological research demonstrates that emotion, as an essential factor in empathy,…

Computation and Language · Computer Science 2024-03-26 Wang Yufeng , Chen Chao , Yang Zhou , Wang Shuhui , Liao Xiangwen

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the target language. Building such a system faces challenges of…

Sound · Computer Science 2023-10-09 Yuke Li , Xinfa Zhu , Yi Lei , Hai Li , Junhui Liu , Danming Xie , Lei Xie

State-of-the-art Text-To-Speech (TTS) models are capable of producing high-quality speech. The generated speech, however, is usually neutral in emotional expression, whereas very often one would want fine-grained emotional control of words…

Sound · Computer Science 2023-03-14 Shijun Wang , Jón Guðnason , Damian Borth

The rapid growth of Speech Emotion Recognition (SER) has diverse global applications, from improving human-computer interactions to aiding mental health diagnostics. However, SER models might contain social bias toward gender, leading to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-06 Yi-Cheng Lin , Haibin Wu , Huang-Cheng Chou , Chi-Chun Lee , Hung-yi Lee

Speech emotion sensing in communication networks has a wide range of applications in real life. In these applications, voice data are transmitted from the user to the central server for storage, processing, and decision making. However,…

Using printed photograph and replaying videos of biometric modalities, such as iris, fingerprint and face, are common attacks to fool the recognition systems for granting access as the genuine user. With the growing online person-to-person…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Joel Stehouwer , Amin Jourabloo , Yaojie Liu , Xiaoming Liu

With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich…

Computation and Language · Computer Science 2025-09-18 Haoyu Wang , Guangyan Zhang , Jiale Chen , Jingyu Li , Yuehai Wang , Yiwen Guo

Emotional support conversation (ESC) helps reduce people's psychological stress and provide emotional value through interactive dialogues. Due to the high cost of crowdsourcing a large ESC corpus, recent attempts use large language models…

Computation and Language · Computer Science 2025-06-23 Zhuang Chen , Yaru Cao , Guanqun Bi , Jincenzi Wu , Jinfeng Zhou , Xiyao Xiao , Si Chen , Hongning Wang , Minlie Huang

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Xiaoxue Gao , Huayun Zhang , Nancy F. Chen