中文
相关论文

相关论文: Representation Loss Minimization with Randomized S…

200 篇论文

In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To…

声音 · 计算机科学 2025-05-27 Jingguang Tian , Xinhui Hu , Xinkang Xu

Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment…

声音 · 计算机科学 2023-01-27 Jianwei Zhang , Julie Liss , Suren Jayasuriya , Visar Berisha

With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution of deepfake…

声音 · 计算机科学 2024-06-11 Yuankun Xie , Ruibo Fu , Zhengqi Wen , Zhiyong Wang , Xiaopeng Wang , Haonnan Cheng , Long Ye , Jianhua Tao

Speech representation models based on the transformer architecture and trained by self-supervised learning have shown great promise for solving tasks such as speech and speaker recognition, keyword spotting, emotion detection, and more.…

计算与语言 · 计算机科学 2024-11-25 Teresa Dorszewski , Lenka Tětková , Lars Kai Hansen

Dimension reduction is often an important step in the analysis of high-dimensional data. PCA is a popular technique to find the best low-dimensional approximation of high-dimensional data. However, classical PCA is very sensitive to…

统计计算 · 统计学 2019-01-14 Holger Cevallos-Valdiviezo , Stefan Van Aelst

Dimension reduction techniques usually lose information in the sense that reconstructed data are not identical to the original data. However, we argue that it is possible to have reconstructed data identically distributed as the original…

机器学习 · 统计学 2026-05-15 Xinwei Shen , Nicolai Meinshausen

Over the recent years, various deep learning-based methods were proposed for extracting a fixed-dimensional embedding vector from speech signals. Although the deep learning-based embedding extraction methods have shown good performance in…

音频与语音处理 · 电气工程与系统科学 2021-12-08 Woo Hyun Kang , Jahangir Alam , Abderrahim Fathan

The high-dimensional data setting, in which p >> n, is a challenging statistical paradigm that appears in many real-world problems. In this setting, learning a compact, low-dimensional representation of the data can substantially help…

机器学习 · 计算机科学 2018-08-07 Micol Marchetti-Bowick , Benjamin J. Lengerich , Ankur P. Parikh , Eric P. Xing

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Unsupervised Domain Adaptation (UDA) addresses the problem of performance degradation due to domain shift between training and testing sets, which is common in computer vision applications. Most existing UDA approaches are based on…

计算机视觉与模式识别 · 计算机科学 2019-05-14 Songsong Wu , Yan Yan , Hao Tang , Jianjun Qian , Jian Zhang , Xiao-Yuan Jing

The slowness principle is a concept inspired by the visual cortex of the brain. It postulates that the underlying generative factors of a quickly varying sensory signal change on a slower time scale. Unsupervised learning of intermediate…

机器学习 · 计算机科学 2021-07-06 Oliver Struckmeier , Kshitij Tiwari , Ville Kyrki

This paper proposes an active learning system for sound event detection (SED). It aims at maximizing the accuracy of a learned SED model with limited annotation effort. The proposed system analyzes an initially unlabeled audio dataset, from…

音频与语音处理 · 电气工程与系统科学 2020-09-10 Shuyang Zhao , Toni Heittola , Tuomas Virtanen

Data quality is a crucial factor in large language models training. While prior work has shown that models trained on smaller, high-quality datasets can outperform those trained on much larger but noisy or low-quality corpora, systematic…

机器学习 · 计算机科学 2026-02-17 Youwei Shu , Shaomian Zheng , Dingnan Jin , Wenjie Qu , Ziyao Guo , Qing Cui , Jun Zhou , Jiaheng Zhang

Digital sound synthesis presents the opportunity to explore vast parameter spaces containing millions of configurations. Quality diversity (QD) evolutionary algorithms offer a promising approach to harness this potential, yet their success…

声音 · 计算机科学 2025-12-03 Björn Þór Jónsson , Çağrı Erdem , Stefano Fasciani , Kyrre Glette

The SOTA in transcription of disfluent and conversational speech has in recent years favored two-stage models, with separate transcription and cleaning stages. We believe that previous attempts at end-to-end disfluency removal have fallen…

音频与语音处理 · 电气工程与系统科学 2023-09-12 Saksham Bassi , Giulio Duregon , Siddhartha Jalagam , David Roth

Stacked denoising autoencoders (SDAs) have been successfully used to learn new representations for domain adaptation. Recently, they have attained record accuracy on standard benchmark tasks of sentiment analysis across different text…

机器学习 · 计算机科学 2012-06-22 Minmin Chen , Zhixiang Xu , Kilian Weinberger , Fei Sha

In this work, we investigate various state-of-the-art (SOTA) speech pre-trained models (PTMs) for their capability to capture prosodic signatures of the generative sources for audio deepfake source attribution (ADSD). These prosodic…

音频与语音处理 · 电气工程与系统科学 2024-12-24 Orchid Chetia Phukan , Drishti Singh , Swarup Ranjan Behera , Arun Balaji Buduru , Rajesh Sharma

Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech as speaker and…

音频与语音处理 · 电气工程与系统科学 2021-12-06 Jie Wang , Jingbei Li , Xintao Zhao , Zhiyong Wu , Shiyin Kang , Helen Meng

The performance of speech and events recognition systems significantly improved recently thanks to deep learning methods. However, some of these tasks remain challenging when algorithms are deployed on robots due to the unseen mechanical…

音频与语音处理 · 电气工程与系统科学 2023-03-08 Pierre-Olivier Lagacé , François Ferland , François Grondin

Dimensionality reduction techniques play important roles in the analysis of big data. Traditional dimensionality reduction approaches, such as principal component analysis (PCA) and linear discriminant analysis (LDA), have been studied…

机器学习 · 计算机科学 2018-05-31 Haozhe Xie , Jie Li , Hanqing Xue