English
Related papers

Related papers: Interpretable Binaural Deep Beamforming Guided by …

200 papers

This paper introduces a new approach to sound source localization using head-related transfer function (HRTF) characteristics, which enable precise full-sphere localization from raw data. While previous research focused primarily on using…

Sound · Computer Science 2024-02-07 Gil Geva , Olivier Warusfel , Shlomo Dubnov , Tammuz Dubnov , Amir Amedi , Yacov Hel-Or

Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack the parallelizability, scalability, and versatility of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-08 Kohei Saijo , Gordon Wichern , François G. Germain , Zexu Pan , Jonathan Le Roux

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Yuzhu Wang , Archontis Politis , Tuomas Virtanen

Beamforming (BF) is essential for enhancing system capacity in fifth generation (5G) and beyond wireless networks, yet exhaustive beam training in ultra-massive multiple-input multiple-output (MIMO) systems incurs substantial overhead. To…

Signal Processing · Electrical Eng. & Systems 2026-02-11 Yanliang Jin , Yunfan Li , Jiang Jun , Yuan Gao , Shengli Liu , Jianbo Du , Zhaohui Yang , Shugong Xu

This paper presents, in the context of multi-channel ASR, a method to adapt a mask based, statistically optimal beamforming approach to a speaker of interest. The beamforming vector of the statistically optimal beamformer is computed by…

Computation and Language · Computer Science 2018-06-21 Tobias Menne , Ralf Schlüter , Hermann Ney

Sensor-aided beamforming reduces the overheads associated with beam training in millimeter-wave (mmWave) multi-input-multi-output (MIMO) communication systems. Most prior work, though, neglects the challenges associated with establishing…

Signal Processing · Electrical Eng. & Systems 2025-09-17 Kartik Patel , Robert W. Heath

Recently, the research on ad-hoc microphone arrays with deep learning has drawn much attention, especially in speech enhancement and separation. Because an ad-hoc microphone array may cover such a large area that multiple speakers may…

Sound · Computer Science 2020-12-02 Ziye Yang , Shanzheng Guan , Xiao-Lei Zhang

This paper proposes a deep speech enhancement method which exploits the high potential of residual connections in a wide neural network architecture, a topology known as Wide Residual Network. This is supported on single dimensional…

Sound · Computer Science 2019-01-04 Dayana Ribas , Jorge Llombart , Antonio Miguel , Luis Vicente

We present a framework that can impose the audio effects and production style from one recording to another by example with the goal of simplifying the audio production process. We train a deep neural network to analyze an input recording…

Sound · Computer Science 2022-07-19 Christian J. Steinmetz , Nicholas J. Bryan , Joshua D. Reiss

Recent applications of pattern recognition techniques on brain connectome classification using functional connectivity (FC) are shifting towards acknowledging the non-Euclidean topology and dynamic aspects of brain connectivity across time.…

Machine Learning · Computer Science 2024-11-12 Sin-Yee Yap , Junn Yong Loo , Chee-Ming Ting , Fuad Noman , Raphael C. -W. Phan , Adeel Razi , David L. Dowe

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

This paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for…

Sound · Computer Science 2020-12-08 Xiaofei Li , Laurent Girin , Fabien Badeig , Radu Horaud

Direct volume rendering (DVR) is a fundamental technique for visualizing volumetric data, where transfer functions (TFs) play a crucial role in extracting meaningful structures. However, designing effective TFs remains unintuitive due to…

Graphics · Computer Science 2025-09-10 Yiyao Wang , Bo Pan , Ke Wang , Han Liu , Jinyuan Mao , Yuxin Liu , Minfeng Zhu , Xiuqi Huang , Weifeng Chen , Bo Zhang , Wei Chen

This paper describes the integration of weighted delay-and-sum beamforming with speech source localization using image processing and robot head visual servoing for source tracking. We take into consideration the fact that the directivity…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-19 José Novoa , Rodrigo Mahu , Alejandro Díaz , Jorge Wuth , Richard Stern , Nestor Becerra Yoma

This paper presents a complete hardware and software pipeline for real-time speech enhancement in noisy and reverberant conditions. The device consists of a microphone array and a camera mounted on eyeglasses, connected to an embedded…

Millimeter wave (mmWave) communication, utilizing beamforming techniques to address the inherent path loss limitation, is considered as one of the key technologies to support ever increasing high throughput and low latency demands of…

Networking and Internet Architecture · Computer Science 2026-02-17 Muhammad Baqer Mollah , Honggang Wang , Mohammad Ataul Karim , Hua Fang

We present an indoor acoustic simulation framework that supports both ultrasonic and audible signaling. The framework opens the opportunity for fast indoor acoustic data generation and positioning development. The improved…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-22 Daan Delabie , Chesney Buyle , Bert Cox , Liesbet Van der Perre , Lieven De Strycker

Recently, a method has been proposed to estimate the direction of arrival (DOA) of a single speaker by minimizing the frequency-averaged Hermitian angle between an estimated relative transfer function (RTF) vector and a database of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-28 Daniel Fejgin , Simon Doclo

To enhance immersive experiences, binaural audio offers spatial awareness of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural…

Sound · Computer Science 2025-06-03 Tianrui Pan , Jie Liu , Zewen Huang , Jie Tang , Gangshan Wu

Telepresence aims to create an immersive but virtual experience of the audio and visual scene at the far end for users at the near end. In this contribution, we propose an array-based binaural rendering system that converts the array…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-07 Yicheng Hsu , Chenghumg Ma , Mingsian R. Bai