中文
相关论文

相关论文: Speaker Placement Agnosticism: Improving the Dista…

200 篇论文

We recently proposed DOVER-Lap, a method for combining overlap-aware speaker diarization system outputs. DOVER-Lap improved upon its predecessor DOVER by using a label mapping method based on globally-informed greedy search. In this paper,…

音频与语音处理 · 电气工程与系统科学 2021-06-07 Desh Raj , Sanjeev Khudanpur

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

声音 · 计算机科学 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

A key task for speech recognition systems is to reduce the mismatch between training and evaluation data that is often attributable to speaker differences. Speaker adaptation techniques play a vital role to reduce the mismatch. Model-based…

声音 · 计算机科学 2024-06-17 Xurong Xie , Xunying Liu , Tan Lee , Lan Wang

Recent works on deep non-linear spatially selective filters demonstrate exceptional enhancement performance with computationally lightweight architectures for stationary speakers of known directions. However, to maintain this performance in…

音频与语音处理 · 电气工程与系统科学 2025-07-08 Jakob Kienegger , Alina Mannanova , Huajian Fang , Timo Gerkmann

In this work we apply deep reinforcement learning to the problems of navigating a three-dimensional environment and inferring the locations of human speaker audio sources within, in the case where the only available information is the raw…

声音 · 计算机科学 2021-11-30 Petros Giannakopoulos , Aggelos Pikrakis , Yannis Cotronis

Background noise reduces speech intelligibility and quality, making speaker verification (SV) in noisy environments a challenging task. To improve the noise robustness of SV systems, additive noise data augmentation method has been commonly…

音频与语音处理 · 电气工程与系统科学 2023-07-21 Wonbin Kim , Hyun-seo Shin , Ju-ho Kim , Jungwoo Heo , Chan-yeong Lim , Ha-Jin Yu

Speaker Identification refers to the process of identifying a person using one's voice from a collection of known speakers. Environmental noise, reverberation and distortion make the task of automatic speaker identification challenging as…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Sabbir Ahmed , Nursadul Mamun , Md Azad Hossain

Speech enhancement in hearing aids remains a difficult task in nonstationary acoustic environments, mainly because current signal processing algorithms rely on fixed, manually tuned parameters that cannot adapt in situ to different users or…

Image-based diagnostic decision support systems (DDSS) utilizing deep learning have the potential to optimize clinical workflows. However, developing DDSS requires extensive datasets with expert annotations and is therefore costly.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Helen Schneider , Sebastian Nowak , Aditya Parikh , Yannik C. Layer , Maike Theis , Wolfgang Block , Alois M. Sprinkart , Ulrike Attenberger , Rafet Sifa

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Classical methods for acoustic scene mapping require the estimation of time difference of arrival (TDOA) between microphones. Unfortunately, TDOA estimation is very sensitive to reverberation and additive noise. We introduce an unsupervised…

音频与语音处理 · 电气工程与系统科学 2024-03-14 Idan Cohen , Ofir Lindenbaum , Sharon Gannot

The problem of mixed signals occurs in many different contexts; one of the most familiar being acoustics. The forward problem in acoustics consists of finding the sound pressure levels at various detectors resulting from sound signals…

数据分析、统计与概率 · 物理学 2007-05-23 Kevin H. Knuth

Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural audio, which mimics human hearing, can capture richer spatial…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Zexu Pan , Shengkui Zhao , Yukun Ma , Haoxu Wang , Yiheng Jiang , Biao Tian , Bin Ma

Photoacoustic imaging (PAI) is an emerging biomedical imaging modality capable of providing both high contrast and high resolution of optical and UltraSound (US) imaging. When a short duration laser pulse illuminates the tissue as a target…

信号处理 · 电气工程与系统科学 2018-01-19 Moein Mozaffarzadeh , Ali Mahloojifar , Mahdi Orooji

Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolution with an enhanced accuracy, a multi-scale approach is a way…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Tae Jin Park , Nithin Rao Koluguri , Jagadeesh Balam , Boris Ginsburg

For computational acoustics, schemes need to have low-dispersion and low-dissipation properties in order to capture the amplitude and phase of the wave correctly. To improve the spectral properties of the scheme, the authors have previously…

计算物理 · 物理学 2021-11-15 Y. H. Li , Y. X. Ren , Y. T. Su

Diffusion magnetic resonance imaging datasets suffer from low Signal-to-Noise Ratio, especially at high b-values. Acquiring data at high b-values contains relevant information and is now of great interest for microstructural and…

计算机视觉与模式识别 · 计算机科学 2016-06-27 Samuel St-Jean , Pierrick Coupé , Maxime Descoteaux

This work considers the problem of locating a single source from noisy range measurements to a set of nodes in a wireless sensor network. We propose two new techniques that we designate as Source Localization with Nuclear Norm (SLNN) and…

最优化与控制 · 数学 2011-11-30 Pınar Oğuz-Ekim , João Gomes , João Xavier , Marko Stošić , Paulo Oliveira

This paper deals with the problem of reconstructing the path of a vehicle in an unknown environment consisting of planar structures using sound. Many systems in the literature do this by using a loudspeaker and microphones mounted on a…

度量几何 · 数学 2024-03-04 Mireille Boutin , Gregor Kemper

We address the problem of online localization and tracking of multiple moving speakers in reverberant environments. The paper has the following contributions. We use the direct-path relative transfer function (DP-RTF), an inter-channel…

声音 · 计算机科学 2019-04-11 Xiaofei Li , Yutong Ban , Laurent Girin , Xavier Alameda-Pineda , Radu Horaud