English
Related papers

Related papers: Improving Speaker Representations Using Contrastiv…

200 papers

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

The use of deep neural networks (DNN) has dramatically elevated the performance of automatic speaker verification (ASV) over the last decade. However, ASV systems can be easily neutralized by spoofing attacks. Therefore, the Spoofing-Aware…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-14 Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin

Many endeavors have sought to develop countermeasure techniques as enhancements on Automatic Speaker Verification (ASV) systems, in order to make them more robust against spoof attacks. As evidenced by the latest ASVspoof 2019…

Sound · Computer Science 2021-09-21 Amir Mohammad Rostami , Mohammad Mehdi Homayounpour , Ahmad Nickabadi

In this paper, we propose a noise robust bottleneck feature representation which is generated by an adversarial network (AN). The AN includes two cascade connected networks, an encoding network (EN) and a discriminative network (DN).…

Sound · Computer Science 2017-06-13 Hong Yu , Zheng-Hua Tan , Zhanyu Ma , Jun Guo

Speaker embeddings achieve promising results on many speaker verification tasks. Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings. In this paper, we introduce phonetic…

Sound · Computer Science 2018-06-15 Yi Liu , Liang He , Jia Liu , Michael T. Johnson

In recent years, using raw waveforms as input for deep networks has been widely explored for the speaker verification system. For example, RawNet and RawNet2 extracted speaker's feature embeddings from waveforms automatically for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-08 Jin Li , Nan Yan , Lan Wang

We propose an end-to-end speaker verification system based on the neural network and trained by a loss function with less computational complexity. The end-to-end speaker verification system in this paper consists of a ResNet architecture…

Sound · Computer Science 2018-09-05 Xuan Shi , Xingjian Du , Mengyao Zhu

Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and…

Machine Learning · Computer Science 2019-11-20 Zhongxin Bai , Xiao-Lei Zhang , Jingdong Chen

In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and text encoder together. Secondly, we show that using this…

Sound · Computer Science 2022-06-30 Andrew Koh , Eng Siong Chng

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

Sound · Computer Science 2021-02-11 Zeqian Li , Jacob Whitehill

The performance of face recognition system degrades when the variability of the acquired faces increases. Prior work alleviates this issue by either monitoring the face quality in pre-processing or predicting the data uncertainty along with…

Computer Vision and Pattern Recognition · Computer Science 2021-07-27 Qiang Meng , Shichao Zhao , Zhida Huang , Feng Zhou

Contrastive learning produces coherent semantic feature embeddings by encouraging positive samples to cluster closely while separating negative samples. However, existing contrastive learning methods lack principled guarantees on coverage…

Machine Learning · Computer Science 2026-03-30 Yahya Alkhatib , Wee Peng Tay

Advances in computer vision and deep learning have blurred the line between deepfakes and authentic media, undermining multimedia credibility through audio-visual forgery. Current multimodal detection methods remain limited by unbalanced…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Zihan Xiong , Xiaohua Wu , Lei Chen , Fangqi Lou

Transformer-based models have shown strong performance in speech deepfake detection, largely due to the effectiveness of the multi-head self-attention (MHSA) mechanism. MHSA provides frame-level attention scores, which are particularly…

Sound · Computer Science 2026-02-05 Tuan Dat Phuong , Duc-Tuan Truong , Long-Vu Hoang , Trang Nguyen Thi Thu

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final…

Sound · Computer Science 2023-06-27 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

Deep learning has dramatically improved the performance of speech recognition systems through learning hierarchies of features optimized for the task at hand. However, true end-to-end learning, where features are learned directly from…

Computation and Language · Computer Science 2016-04-06 Zhenyao Zhu , Jesse H. Engel , Awni Hannun

Feature extraction plays an important role as a front-end processing block in speaker identification (SI) process. Most of the SI systems utilize like Mel-Frequency Cepstral Coefficients (MFCC), Perceptual Linear Prediction (PLP), Linear…

Sound · Computer Science 2015-03-19 Md. Sahidullah , Sandipan Chakroborty , Goutam Saha

In far-field speaker verification, the performance of speaker embeddings is susceptible to degradation when there is a mismatch between the conditions of enrollment and test speech. To solve this problem, we propose the feature-level and…

Sound · Computer Science 2021-06-18 Li Zhang , Qing Wang , Kong Aik Lee , Lei Xie , Haizhou Li

A good training set for speech spoofing countermeasures requires diverse TTS and VC spoofing attacks, but generating TTS and VC spoofed trials for a target speaker may be technically demanding. Instead of using full-fledged TTS and VC…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Xin Wang , Junichi Yamagishi

Automatic speaker verification systems are vulnerable to a variety of access threats, prompting research into the formulation of effective spoofing detection systems to act as a gate to filter out such spoofing attacks. This study…

Sound · Computer Science 2022-11-21 Zhenyu Wang , John H. L. Hansen
‹ Prev 1 4 5 6 7 8 10 Next ›