中文
相关论文

相关论文: Reverberant Sound Localization with a Robot Head B…

200 篇论文

This article is a survey on deep learning methods for single and multiple sound source localization. We are particularly interested in sound source localization in indoor/domestic environment, where reverberation and diffuse noise are…

声音 · 计算机科学 2022-07-20 Pierre-Amaury Grumiaux , Srđan Kitić , Laurent Girin , Alexandre Guérin

Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embeddings using deep neural networks for SV systems has gone…

声音 · 计算机科学 2022-05-27 Nan Zhang , Jianzong Wang , Zhenhou Hong , Chendong Zhao , Xiaoyang Qu , Jing Xiao

The Radiative transfer coherent backscattering (RT-CB) code is extended to apply to dense discrete random media of optically soft spherical particles. This is achieved by utilizing the well-known static-structure-factor (SSF) correction…

材料科学 · 物理学 2023-11-23 Johannes Markkanen , Antti Penttilä

Sound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on utilizing deep…

Estimating the position of a speech source based on time-differences-of-arrival (TDOAs) is often adversely affected by background noise and reverberation. A popular method to estimate the TDOA between a microphone pair involves maximizing a…

音频与语音处理 · 电气工程与系统科学 2025-07-28 Klaus Brümann , Kouei Yamaoka , Nobutaka Ono , Simon Doclo

Acoustic reflector localization is an important issue in audio signal processing, with direct applications in spatial audio, scene reconstruction, and source separation. Several methods have recently been proposed to estimate the 3D…

声音 · 计算机科学 2017-01-06 Luca Remaggi , Philip J. B. Jackson , Philip Coleman , Wenwu Wang

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this…

Audio anti-spoofing systems are typically formulated as binary classifiers distinguishing bona fide from spoofed speech. This assumption fails under layered generative processing, where benign transformations introduce distributional shifts…

声音 · 计算机科学 2026-03-17 Shree Harsha Bokkahalli Satish , Harm Lameris , Joakim Gustafson , Éva Székely

Deep learning models are widely applied in the signal processing community, yet their inner working procedure is often treated as a black box. In this paper, we investigate the use of eXplainable Artificial Intelligence (XAI) techniques to…

音频与语音处理 · 电气工程与系统科学 2024-04-29 Luca Comanducci , Fabio Antonacci , Augusto Sarti

This paper proposes a practical approach to estimate the direct-to-reverberant energy ratio (DRR) using a spherical microphone array without having knowledge of the source signal. We base our estimation on a theoretical relationship between…

声音 · 计算机科学 2015-11-02 Hanchi Chen , Prasanga N. Samarasinghe , Thushara D. Abhayapala , Wen Zhang

Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it…

机器人学 · 计算机科学 2024-09-11 Yue Li , Koen V. Hindriks , Florian A. Kunneman

In this paper, we propose an efficient technique for estimating individual power spectral density (PSD) components, i.e., PSD of each desired sound source as well as of noise and reverberation, in a multi-source reverberant sound scene with…

声音 · 计算机科学 2018-05-17 Abdullah Fahim , Prasanga N. Samarasinghe , Thushara D. Abhayapala

Automatic speech recognition (ASR) on multi-talker recordings is challenging. Current methods using 3D spatial data from multi-channel audio and visual cues focus mainly on direct waves from the target speaker, overlooking reflection wave…

音频与语音处理 · 电气工程与系统科学 2024-06-13 Yiwen Shao , Shi-Xiong Zhang , Dong Yu

Automatic speech recognition (ASR) has recently become an important challenge when using deep learning (DL). It requires large-scale training datasets and high computational and storage resources. Moreover, DL techniques and machine…

声音 · 计算机科学 2023-08-01 Hamza Kheddar , Yassine Himeur , Somaya Al-Maadeed , Abbes Amira , Faycal Bensaali

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is…

计算与语言 · 计算机科学 2024-06-14 Amit Meghanani , Thomas Hain

Speech dereverberation in distant-microphone scenarios remains challenging due to the high correlation between reverberation and target signals, often leading to poor generalization in real-world environments. We propose IF-CorrNet, a…

音频与语音处理 · 电气工程与系统科学 2026-03-17 Ui-Hyeop Shin , Jun Hyung Kim , Jangyeon Kim , Wooseok Kim , Hyung-Min Park

Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose ReverbMiipher, an SR model extending parametric resynthesis…

We present an extension of the linear sampling method for solving the sound-soft inverse acoustic scattering problem with randomly distributed point sources. The theoretical justification of our sampling method is based on the…

数值分析 · 数学 2023-03-21 Josselin Garnier , Houssem Haddar , Hadrien Montanelli

Deep learning is an emerging technology that is considered one of the most promising directions for reaching higher levels of artificial intelligence. Among the other achievements, building computers that understand speech represents a…

计算与语言 · 计算机科学 2017-12-19 Mirco Ravanelli

In recent years, Rectified flow (RF) has gained considerable popularity largely due to its generation efficiency and state-of-the-art performance. In this paper, we investigate the degree to which RF automatically adapts to the intrinsic…

机器学习 · 统计学 2026-02-24 Saptarshi Roy , Alessandro Rinaldo , Purnamrita Sarkar