中文
相关论文

相关论文: Xi+: Uncertainty Supervision for Robust Speaker Em…

200 篇论文

In this paper, we propose a simple but powerful unsupervised learning method for speaker recognition, namely Contrastive Equilibrium Learning (CEL), which increases the uncertainty on nuisance factors latent in the embeddings by employing…

音频与语音处理 · 电气工程与系统科学 2020-10-23 Sung Hwan Mun , Woo Hyun Kang , Min Hyun Han , Nam Soo Kim

Uncertainty estimation is important for ensuring safety and robustness of AI systems. While most research in the area has focused on un-structured prediction tasks, limited work has investigated general uncertainty estimation approaches for…

机器学习 · 统计学 2021-02-12 Andrey Malinin , Mark Gales

Transformer models using segment-based processing have been an effective architecture for simultaneous speech translation. However, such models create a context mismatch between training and inference environments, hindering potential…

计算与语言 · 计算机科学 2023-07-06 Matthew Raffel , Drew Penney , Lizhong Chen

The accuracy of automated speaker recognition is negatively impacted by change in emotions in a person's speech. In this paper, we hypothesize that speaker identity is composed of various vocal style factors that may be learned from…

音频与语音处理 · 电气工程与系统科学 2023-08-04 Morgan Sandler , Arun Ross

Transformer-based language models have set new benchmarks across a wide range of NLP tasks, yet reliably estimating the uncertainty of their predictions remains a significant challenge. Existing uncertainty estimation (UE) techniques often…

机器学习 · 计算机科学 2024-09-18 Elizaveta Kostenok , Daniil Cherniavskii , Alexey Zaytsev

Over the recent years, various deep learning-based embedding methods have been proposed and have shown impressive performance in speaker verification. However, as in most of the classical embedding techniques, the deep learning-based…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Woo Hyun Kang , Sung Hwan Mun , Min Hyun Han , Nam Soo Kim

Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue,…

声音 · 计算机科学 2025-08-19 Hyebin Ahn , Kangwook Jang , Hoirin Kim

In recent years, self-supervised learning paradigm has received extensive attention due to its great success in various down-stream tasks. However, the fine-tuning strategies for adapting those pre-trained models to speaker verification…

音频与语音处理 · 电气工程与系统科学 2022-10-05 Junyi Peng , Oldrich Plchot , Themos Stafylakis , Ladislav Mosner , Lukas Burget , Jan Cernocky

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

音频与语音处理 · 电气工程与系统科学 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

Video activity anticipation aims to predict what will happen in the future, embracing a broad application prospect ranging from robot vision and autonomous driving. Despite the recent progress, the data uncertainty issue, reflected as the…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Zhaobo Qi , Shuhui Wang , Weigang Zhang , Qingming Huang

The performance of deep learning models depends significantly on their capacity to encode input features efficiently and decode them into meaningful outputs. Better input and output representation has the potential to boost models'…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Ahmed Adel Attia , Yashish M. Siriwardena , Carol Espy-Wilson

In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To…

声音 · 计算机科学 2025-05-27 Jingguang Tian , Xinhui Hu , Xinkang Xu

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority…

音频与语音处理 · 电气工程与系统科学 2022-02-22 Viet Anh Trinh , Sebastian Braun

Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under…

音频与语音处理 · 电气工程与系统科学 2024-02-06 Kerlos Atia Abdalmalak , Ascensión Gallardo-Antol'in

Human emotion understanding is pivotal in making conversational technology mainstream. We view speech emotion understanding as a perception task which is a more realistic setting. With varying contexts (languages, demographics, etc.)…

人工智能 · 计算机科学 2023-09-29 Payal Mohapatra , Akash Pandey , Yueyuan Sui , Qi Zhu

The explosion of available speech data and new speaker modeling methods based on deep neural networks (DNN) have given the ability to develop more robust speaker recognition systems. Among DNN speaker modelling techniques, x-vector system…

声音 · 计算机科学 2020-06-30 Mohammad Mohammadamini , Driss Matrouf

This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture…

Unsupervised representation learning has shown remarkable achievement by reducing the performance gap with supervised feature learning, especially in the image domain. In this study, to extend the technique of unsupervised learning to the…

音频与语音处理 · 电气工程与系统科学 2020-10-23 Jangho Lee , Jaihyun Koh , Sungroh Yoon

Acoustic models in real-time speech recognition systems typically stack multiple unidirectional LSTM layers to process the acoustic frames over time. Performance improvements over vanilla LSTM architectures have been reported by prepending…

音频与语音处理 · 电气工程与系统科学 2020-07-02 Maarten Van Segbroeck , Harish Mallidih , Brian King , I-Fan Chen , Gurpreet Chadha , Roland Maas

Machine learning techniques are an active area of research for speech enhancement for hearing aids, with one particular focus on improving the intelligibility of a noisy speech signal. Recent work has shown that feature encodings from…

声音 · 计算机科学 2024-07-19 Robert Sutherland , George Close , Thomas Hain , Stefan Goetze , Jon Barker