English
Related papers

Related papers: DSNet: Disentangled Siamese Network with Neutral C…

200 papers

Speaker clustering is the task of identifying the unique speakers in a set of audio recordings (each belonging to exactly one speaker) without knowing who and how many speakers are present in the entire data, which is essential for speaker…

Sound · Computer Science 2025-09-30 Chaohao Lin , Xu Zheng , Kaida Wu , Peihao Xiang , Ou Bai

Recent advances in wireless communication with the enormous demands of sensing ability have given rise to the integrated sensing and communication (ISAC) technology, among which passive sensing plays an important role. The main challenge of…

Information Theory · Computer Science 2023-07-31 Wangjun Jiang , Dingyou Ma , Zhiqing Wei , Zhiyong Feng , Ping Zhang

Advent of modern deep learning techniques has given rise to advancements in the field of Speech Emotion Recognition (SER). However, most systems prevalent in the field fail to generalize to speakers not seen during training. This study…

Computation and Language · Computer Science 2024-06-21 Arnav Goel , Medha Hira , Anubha Gupta

Recent advancements in speaker verification techniques show promise, but their performance often deteriorates significantly in challenging acoustic environments. Although speech enhancement methods can improve perceived audio quality, they…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Adam Katav , Yair Moshe , Israel Cohen

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts…

Machine Learning · Computer Science 2022-11-08 Itai Gat , Hagai Aronowitz , Weizhong Zhu , Edmilson Morais , Ron Hoory

Detecting emotions directly from a speech signal plays an important role in effective human-computer interactions. Existing speech emotion recognition models require massive computational and storage resources, making them hard to implement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-08 Arya Aftab , Alireza Morsali , Shahrokh Ghaemmaghami , Benoit Champagne

Deep convolutional neural networks (DCNNs) based remote sensing (RS) image semantic segmentation technology has achieved great success used in many real-world applications such as geographic element analysis. However, strong dependency on…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Qi Zhao , Shuchang Lyu , Binghao Liu , Lijiang Chen , Hongbo Zhao

Channel state information (CSI) feedback is critical for achieving the promised advantages of enhancing spectral and energy efficiencies in massive multiple-input multiple-output (MIMO) wireless communication systems. Deep learning…

Information Theory · Computer Science 2024-03-29 Suhang Fan , Wei Xu , Renjie Xie , Shi Jin , Derrick Wing Kwan Ng , Naofal Al-Dhahir

This work presents a framework based on feature disentanglement to learn speaker embeddings that are robust to environmental variations. Our framework utilises an auto-encoder as a disentangler, dividing the input speaker embedding into…

Sound · Computer Science 2024-06-21 KiHyun Nam , Hee-Soo Heo , Jee-weon Jung , Joon Son Chung

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Youzhi Tu , Man-Wai Mak , Jen-Tzung Chien

Self-supervised learning in speech involves training a speech representation network on a large-scale unannotated speech corpus, and then applying the learned representations to downstream tasks. Since the majority of the downstream tasks…

In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a speech utterance into features such as linguistic content,…

Sound · Computer Science 2023-02-17 Shehzeen Hussain , Paarth Neekhara , Jocelyn Huang , Jason Li , Boris Ginsburg

Deep neural networks are susceptible to learn biased models with entangled feature representations, which may lead to subpar performances on various downstream tasks. This is particularly true for under-represented classes, where a lack of…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Sanghyeok Chu , Dongwan Kim , Bohyung Han

Speaker verification, as a biometric authentication mechanism, has been widely used due to the pervasiveness of voice control on smart devices. However, the task of "in-the-wild" speaker verification is still challenging, considering the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Jianwei Tai , Xiaoqi Jia , Qingjia Huang , Weijuan Zhang , Haichao Du , Shengzhi Zhang

Alignment of contrast and non-contrast-enhanced imaging is essential for the quantification of changes in several biomedical applications. In particular, the extraction of cartilage shape from contrast-enhanced Computed Tomography (CT) of…

Image and Video Processing · Electrical Eng. & Systems 2021-01-26 Jian-Qing Zheng , Ngee Han Lim , Bartlomiej W. Papiez

Speech super-resolution (SSR) aims to predict a high resolution (HR) speech signal from its low resolution (LR) corresponding part. Most neural SSR models focus on producing the final result in a noise-free environment by recovering the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-11 Junkang Yang , Hongqing Liu , Lu Gan , Yi Zhou

EEG-based tinnitus classification is a valuable tool for tinnitus diagnosis, research, and treatments. Most current works are limited to a single dataset where data patterns are similar. But EEG signals are highly non-stationary, resulting…

Signal Processing · Electrical Eng. & Systems 2022-11-08 Yun Li , Zhe Liu , Lina Yao , Jessica J. M. Monaghan , David McAlpine

Sarcasm is a nuanced and often misinterpreted form of communication, especially in text, where tone and body language are absent. This paper proposes a modular deep learning framework for sarcasm detection, leveraging Deep Convolutional…

Computation and Language · Computer Science 2025-10-14 Manas Zambre , Sarika Bobade

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

Computer Vision and Pattern Recognition · Computer Science 2019-10-03 Gaurav Mittal , Baoyuan Wang