中文
相关论文

相关论文: I'm Sorry for Your Loss: Spectrally-Based Audio Di…

200 篇论文

Existing metric learning losses can be categorized into two classes: pair-based and proxy-based losses. The former class can leverage fine-grained semantic relations between data points, but slows convergence in general due to its high…

计算机视觉与模式识别 · 计算机科学 2020-04-01 Sungyeon Kim , Dongwon Kim , Minsu Cho , Suha Kwak

This paper presents NOMAD (Non-Matching Audio Distance), a differentiable perceptual similarity metric that measures the distance of a degraded signal against non-matching references. The proposed method is based on learning deep feature…

声音 · 计算机科学 2024-01-22 Alessandro Ragano , Jan Skoglund , Andrew Hines

We propose the novel task of distance-based sound separation, where sounds are separated based only on their distance from a single microphone. In the context of assisted listening devices, proximity provides a simple criterion for sound…

声音 · 计算机科学 2022-07-04 Katharine Patterson , Kevin Wilson , Scott Wisdom , John R. Hershey

An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our main proposition is that it is harder to learn sounds in…

声音 · 计算机科学 2020-07-02 Anurag Kumar , Vamsi Krishna Ithapu

Recent works on speech spoofing countermeasures still lack generalization ability to unseen spoofing attacks. This is one of the key issues of ASVspoof challenges especially with the rapid development of diverse and high-quality spoofing…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Monisankha Pal , Aditya Raikar , Ashish Panda , Sunil Kumar Kopparapu

Early visual to auditory substitution devices encode 2D monocular images into sounds while more recent devices use distance information from 3D sensors. This study assesses whether the addition of sound-encoded distance in recent systems…

人机交互 · 计算机科学 2022-04-14 Louis Commère , Jean Rouat

This work introduces a feature extracted from stereophonic/binaural audio signals aiming to represent a measure of perceived quality degradation in processed spatial auditory scenes. The feature extraction technique is based on a simplified…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Pablo M. Delgado , Jürgen Herre

Automated dysarthria detection and severity assessment from speech have attracted significant research attention due to their potential clinical impact. Despite rapid progress in acoustic modeling and deep learning, models still fall short…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Krishna Gurugubelli

We study the effect of memory on synchronization of identical chaotic systems driven by common external noises. Our examples show that while in general synchronization transition becomes more difficult to meet when memory range increases,…

混沌动力学 · 物理学 2009-11-11 Rafael Morgado , Michal Ciesla , Lech Longa , Fernando A. Oliveira

This paper presents a simple yet effective method to achieve prosody transfer from a reference speech signal to synthesized speech. The main idea is to incorporate well-known acoustic correlates of prosody such as pitch and loudness…

声音 · 计算机科学 2020-05-19 Siddharth Gururani , Kilol Gupta , Dhaval Shah , Zahra Shakeri , Jervis Pinto

We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations at the required temporal and frequency resolution is a…

音频与语音处理 · 电气工程与系统科学 2020-09-07 Beat Gfeller , Christian Frank , Dominik Roblek , Matt Sharifi , Marco Tagliasacchi , Mihajlo Velimirović

Current text-to-speech algorithms produce realistic fakes of human voices, making deepfake detection a much-needed area of research. While researchers have presented various techniques for detecting audio spoofs, it is often unclear exactly…

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between the respective spectrograms of the singing and…

声音 · 计算机科学 2020-02-25 Sanna Wager , George Tzanetakis , Cheng-i Wang , Minje Kim

Deep learning-based speech enhancement for real-time applications recently made large advancements. Due to the lack of a tractable perceptual optimization target, many myths around training losses emerged, whereas the contribution to…

音频与语音处理 · 电气工程与系统科学 2020-09-28 Sebastian Braun , Ivan Tashev

Pitch estimation is to estimate the fundamental frequency and the midi number and plays a critical role in music signal analysis and vocal signal processing. In this work, we proposed a new architecture based on a learning-based enhancement…

声音 · 计算机科学 2023-05-09 Yu Cheng Hung , Ping Hung Chen , Jian Jiun Ding

Objective evaluation of synthetic speech quality remains a critical challenge. Human listening tests are the gold standard, but costly and impractical at scale. Fr\'echet Distance has emerged as a promising alternative, yet its reliability…

声音 · 计算机科学 2026-01-30 June-Woo Kim , Dhruv Agarwal , Federica Cerina

Acoustical mismatch among training and testing phases degrades outstandingly speech recognition results. This problem has limited the development of real-world nonspecific applications, as testing conditions are highly variant or even…

声音 · 计算机科学 2013-05-13 Rashmi Makhijani , Urmila Shrawankar , V M Thakare

We investigated, by using auditory models, how three perceptual parameters, loudness, pitch and sharpness, determine human echolocation. We used acoustic recordings from two previous studies, both from stationary situations, and their…

神经元与认知 · 定量生物学 2018-01-31 Bo N. Schenkman , Vijay Kiran Gidla

Texture synthesis techniques based on matching the Gram matrix of feature activations in neural networks have achieved spectacular success in the image domain. In this paper we extend these techniques to the audio domain. We demonstrate…

声音 · 计算机科学 2018-06-22 Joseph Antognini , Matt Hoffman , Ron J. Weiss

This article studies the effects of inter-channel time and level differences in stereophonic reproduction on perceived localization uncertainty, which is defined as how difficult it is for a listener to tell where a sound source is located.…

音频与语音处理 · 电气工程与系统科学 2020-09-08 Enzo De Sena , Zoran Cvetkovic , Huseyin Hacihabiboglu , Marc Moonen , Toon van Waterschoot