中文
相关论文

相关论文: I'm Sorry for Your Loss: Spectrally-Based Audio Di…

200 篇论文

Consumer speech recognition systems do not work as well for many people with speech diferences, such as stuttering, relative to the rest of the general population. However, what is not clear is the degree to which these systems do not work,…

The ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with…

音频与语音处理 · 电气工程与系统科学 2021-12-15 Gabriel Mittag , Saman Zadtootaghaj , Thilo Michael , Babak Naderi , Sebastian Möller

There is extensive interest in metric learning methods for image retrieval. Many metric learning loss functions focus on learning a correct ranking of training samples, but strongly overfit semantically inconsistent labels and require a…

机器学习 · 计算机科学 2023-06-05 Christopher Liao , Theodoros Tsiligkaridis , Brian Kulis

Supervised machine learning utilizes large datasets, often with ground truth labels annotated by humans. While some data points are easy to classify, others are hard to classify, which reduces the inter-annotator agreement. This causes…

人机交互 · 计算机科学 2023-02-14 Andrea Papenmeier , Dagmar Kern , Daniel Hienert , Yvonne Kammerer , Christin Seifert

The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary…

声音 · 计算机科学 2025-09-10 Bin Hu , Kunyang Huang , Daehan Kwak , Meng Xu , Kuan Huang

Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic evaluation metrics and…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Huan Zhang , Jinhua Liang , Huy Phan , Wenwu Wang , Emmanouil Benetos

As spatial audio is enjoying a surge in popularity, data-driven machine learning techniques that have been proven successful in other domains are increasingly used to process head-related transfer function measurements. However, these…

音频与语音处理 · 电气工程与系统科学 2022-12-09 Johan Pauwels , Lorenzo Picinali

Generative models have gained significant attention for their ability to produce realistic synthetic data that supplements the quantity of real-world datasets. While recent studies show performance improvements in wireless sensing tasks by…

机器学习 · 计算机科学 2025-07-01 Chen Gong , Bo Liang , Wei Gao , Chenren Xu

We propose an algorithm for the blind separation of single-channel audio signals. It is based on a parametric model that describes the spectral properties of the sounds of musical instruments independently of pitch. We develop a novel…

音频与语音处理 · 电气工程与系统科学 2021-02-03 Sören Schulze , Emily J. King

Time-reversal symmetry breaking is a key feature of nearly all natural sounds, caused by the physics of sound production. While attention has been paid to the response of the auditory system to "natural stimuli," very few psychophysical…

神经元与认知 · 定量生物学 2013-01-04 Jacob N. Oppenheim , Pavel Isakov , Marcelo O. Magnasco

Accurate stereo depth estimation plays a critical role in various 3D tasks in both indoor and outdoor environments. Recently, learning-based multi-view stereo methods have demonstrated competitive performance with a limited number of views.…

计算机视觉与模式识别 · 计算机科学 2020-06-02 Uday Kusupati , Shuo Cheng , Rui Chen , Hao Su

The goal of this contribution is to use a parametric speech synthesis system for reducing background noise and other interferences from recorded speech signals. In a first step, Hidden Markov Models of the synthesis system are trained. Two…

声音 · 计算机科学 2017-07-06 Daniel Dzibela , Armin Sehr

Despite speech recognition systems achieving low word error rates on standard benchmarks, they often fail on short, high-stakes utterances in real-world deployments. Here, we study this failure mode in a high-stakes task: the transcription…

人工智能 · 计算机科学 2026-02-18 Kaitlyn Zhou , Martijn Bartelds , Federico Bianchi , James Zou

Medical audio classification remains challenging due to low signal-to-noise ratios, subtle discriminative features, and substantial intra-class variability, often compounded by class imbalance and limited training data. Synthetic data…

声音 · 计算机科学 2026-02-04 David McShannon , Anthony Mella , Nicholas Dietrich

Do object part localization methods produce bilaterally symmetric results on mirror images? Surprisingly not, even though state of the art methods augment the training set with mirrored images. In this paper we take a closer look into this…

计算机视觉与模式识别 · 计算机科学 2015-01-22 Heng Yang , Ioannis Patras

The acoustic cues used by humans and other animals to localise sounds are subtle, and change during and after development. This means that we need to constantly relearn or recalibrate the auditory spatial map throughout our lifetimes. This…

神经与进化计算 · 计算机科学 2025-04-18 Yang Chu , Wayne Luk , Dan Goodman

The effect of hearing impairment on speech perception was described by Plomp (1978) as a sum of a loss of class A, due to signal attenuation, and a loss of class D, due to signal distortion. While a loss of class A can be compensated by…

音频与语音处理 · 电气工程与系统科学 2021-02-25 Marc René Schädler

Detecting and mitigating bias in speaker verification systems is important, as datasets, processing choices and algorithms can lead to performance differences that systematically favour some groups of people while disadvantaging others.…

音频与语音处理 · 电气工程与系统科学 2024-08-27 Wiebke Hutiri , Tanvina Patel , Aaron Yi Ding , Odette Scharenborg

The NLP community has broadly focused on text-only approaches of cognitive state tasks, but audio can provide vital missing cues through prosody. We posit that text-to-speech models learn to track aspects of cognitive state in order to…

声音 · 计算机科学 2025-02-12 Adil Soubki , John Murzaku , Peter Zeng , Owen Rambow

Audio fingerprinting (AFP) allows the identification of unknown audio content by extracting compact representations, termed audio fingerprints, that are designed to remain robust against common audio degradations. Neural AFP methods often…