English
Related papers

Related papers: Can Emotion Fool Anti-spoofing?

200 papers

With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple…

Multimedia · Computer Science 2022-09-19 Davide Salvi , Brian Hosler , Paolo Bestagini , Matthew C. Stamm , Stefano Tubaro

With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financial gain) has brought huge hazards in our daily lives.…

Sound · Computer Science 2024-10-29 Zhisheng Zhang , Qianyi Yang , Derui Wang , Pengyang Huang , Yuxin Cao , Kai Ye , Jie Hao

Emotional intelligence in conversational AI is crucial across domains like human-computer interaction. While numerous models have been developed, they often overlook the complexity and ambiguity inherent in human emotions. In the era of…

Sound · Computer Science 2025-05-27 Jule Valendo Halim , Siyi Wang , Hong Jia , Ting Dang

Emotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech…

Computation and Language · Computer Science 2023-01-02 Kun Zhou , Berrak Sisman , Rajib Rana , B. W. Schuller , Haizhou Li

Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided…

Computation and Language · Computer Science 2026-01-13 Jiaqi Qiao , Xiujuan Xu , Xinran Li , Yu Liu

Electroencephalographic (EEG) signals have long been applied in the field of affective brain-computer interfaces (aBCIs). Cross-subject EEG-based emotion recognition has demonstrated significant potential in practical applications due to…

Machine Learning · Computer Science 2025-12-23 Yici Liu , Qi Wei Oung , Hoi Leong Lee

Voiced Electromyography (EMG)-to-Speech (V-ETS) models reconstruct speech from muscle activity signals, facilitating applications such as neurolaryngologic diagnostics. Despite its potential, the advancement of V-ETS is hindered by a…

Sound · Computer Science 2026-01-13 Xiaodan Chen , Xiaoxue Gao , Mathias Quoy , Alexandre Pitti , Nancy F. Chen

The automatic speaker verification spoofing (ASVspoof) challenge series is crucial for enhancing the spoofing consideration and the countermeasures growth. Although the recent ASVspoof 2019 validation results indicate the significant…

Sound · Computer Science 2022-09-27 Chenlei Hu , Ruohua Zhou

During the last few years, spoken language technologies have known a big improvement thanks to Deep Learning. However Deep Learning-based algorithms require amounts of data that are often difficult and costly to gather. Particularly,…

Sound · Computer Science 2019-01-15 Noé Tits , Kevin El Haddad , Thierry Dutoit

While multiple emotional speech corpora exist for commonly spoken languages, there is a lack of functional datasets for smaller (spoken) languages, such as Danish. To our knowledge, Danish Emotional Speech (DES), published in 1997, is the…

Computation and Language · Computer Science 2025-08-21 Maja J. Hjuler , Harald V. Skat-Rørdam , Line H. Clemmensen , Sneha Das

Understanding the reason for emotional support response is crucial for establishing connections between users and emotional support dialogue systems. Previous works mostly focus on generating better responses but ignore interpretability,…

Computation and Language · Computer Science 2024-06-18 Tenggan Zhang , Xinjie Zhang , Jinming Zhao , Li Zhou , Qin Jin

A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for…

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their…

Sound · Computer Science 2025-04-17 Tian-Hao Zhang , Jiawei Zhang , Jun Wang , Xinyuan Qian , Xu-Cheng Yin

Detecting spoofing attempts of automatic speaker verification (ASV) systems is challenging, especially when using only one modeling approach. For robustness, we use both deep neural networks and traditional machine learning models and…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-05 Bhusan Chettri , Daniel Stoller , Veronica Morfi , Marco A. Martínez Ramírez , Emmanouil Benetos , Bob L. Sturm

Sentiment analysis has various application scenarios in software engineering (SE), such as detecting developers' emotions in commit messages and identifying their opinions on Q&A forums. However, commonly used out-of-the-box sentiment…

Software Engineering · Computer Science 2019-07-05 Zhenpeng Chen , Yanbin Cao , Xuan Lu , Qiaozhu Mei , Xuanzhe Liu

Understanding individual, group and event level emotions along with contextual information is crucial for analyzing a multi-person social situation. To achieve this, we frame emotion comprehension as the task of predicting fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Anubhav Kataria , Surbhi Madan , Shreya Ghosh , Tom Gedeon , Abhinav Dhall

This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotional speech synthesis often needs manual labels or reference…

Sound · Computer Science 2020-11-18 Yi Lei , Shan Yang , Lei Xie

Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are aggregated into a single representation via a lightweight…

Sound · Computer Science 2026-02-02 Seungu Han , Sungho Lee , Kyogu Lee

Speaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker introduces…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Ze Li , Yao Shi , Yunfei Xu , Ming Li

The current speech anti-spoofing countermeasures (CMs) show excellent performance on specific datasets. However, removing the silence of test speech through Voice Activity Detection (VAD) can severely degrade performance. In this paper, the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-22 Yuxiang Zhang , Zhuo Li , Jingze Lu , Hua Hua , Wenchao Wang , Pengyuan Zhang
‹ Prev 1 8 9 10 Next ›