English
Related papers

Related papers: Weighted delay-and-sum beamforming guided by visua…

200 papers

Sound source tracking is commonly performed using classical array-processing algorithms, while machine-learning approaches typically rely on precise source position labels that are expensive or impractical to obtain. This paper introduces a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-12 Luan Vinícius Fiorio , Ivana Nikoloska , Bruno Defraene , Alex Young , Johan David , Ronald M. Aarts

Recent research advances in deep neural network (DNN)-based beamformers have shown great promise for speech enhancement under adverse acoustic conditions. Different network architectures and input features have been explored in estimating…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-24 Hsinyu Chang , Yicheng Hsu , Mingsian R. Bai

In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-18 Ilai Zaidel , Sharon Gannot

High-precision tiny object alignment remains a common and critical challenge for humanoid robots in real-world. To address this problem, this paper proposes a vision-based framework for precisely estimating and controlling the relative…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Jialong Xue , Wei Gao , Yu Wang , Chao Ji , Dongdong Zhao , Shi Yan , Shiwu Zhang

The cooperation of a pair of robot manipulators is required to manipulate a target object without any fixtures. The conventional control methods coordinate the end-effector pose of each manipulator with that of the other using their…

Robotics · Computer Science 2025-10-08 Zizhe Zhang , Yuan Yang , Wenqiang Zuo , Guangming Song , Aiguo Song , Yang Shi

Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually…

Sound · Computer Science 2024-06-06 Theo Mariotte , Anthony Larcher , Silvio Montresor , Jean-Hugh Thomas

Sound sources localization using multichannel signal processing has been a subject of active research for decades. In recent years, the use of deep learning in audio signal processing has allowed to drastically improve performances for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Hadrien Pujol , Éric Bavu , Alexandre Garcia

Visual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally applicable without task-specific training, they often…

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

This paper investigates the performance of Binaural Signal Matching (BSM) methods for near-field sound reproduction using a wearable glasses-mounted microphone array. BSM is a flexible, signal-independent approach for binaural rendering…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Sapir Goldring , Zamir Ben Hur , David Lou Alon , Chad McKell , Sebastian Prepelita , Boaz Rafaely

The spatial covariance matrix has been considered to be significant for beamformers. Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and…

Sound · Computer Science 2021-09-03 Andong Li , Wenzhe Liu , Chengshi Zheng , Xiaodong Li

The pervasive nature of wireless telecommunication has made it the foundation for mainstream technologies like automation, smart vehicles, virtual reality, and unmanned aerial vehicles. As these technologies experience widespread adoption…

Information Theory · Computer Science 2023-07-21 Jaspreet Kaur , Satyam Bhatti , Olaoluwa R Popoola , Muhammad Ali Imran , Rami Ghannam , Qammer H Abbasi , Hasan T Abbas

With explosively increasing demands for unmanned aerial vehicle (UAV) applications, reliable link acquisition for serving UAVs is required. Considering the dynamic characteristics of UAV, it is hugely challenging to persist a reliable link…

Signal Processing · Electrical Eng. & Systems 2020-11-23 Ha-Lim Song , Young-Chai Ko

Millimeter-wave (mmWave) and terahertz (THz) communications require beamforming to acquire adequate receive signal-to-noise ratio (SNR). To find the optimal beam, current beam management solutions perform beam training over a large number…

Signal Processing · Electrical Eng. & Systems 2021-11-30 Shuaifeng Jiang , Ahmed Alkhateeb

This paper proposes a method for estimating a convolutional beamformer that can perform denoising and dereverberation simultaneously in an optimal way. The application of dereverberation based on a weighted prediction error (WPE) method…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-07 Tomohiro Nakatani , Keisuke Kinoshita

This article discusses aeroacoustic imaging methods based on correlation measurements in the frequency domain. Standard methods in this field assume that the estimated correlation matrix is superimposed with additive white noise. In this…

Signal Processing · Electrical Eng. & Systems 2020-12-30 Hans-Georg Raumer , Carsten Spehr , Thorsten Hohage , Daniel Ernst

Interfering sources, background noise and reverberation degrade speech quality and intelligibility in hearing aid applications. In this paper, we present an adaptive algorithm aiming at dereverberation, noise and interferer reduction and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Henri Gode , Simon Doclo

Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Francois Grondin , Jean-Samuel Lauzon , Jonathan Vincent , Francois Michaud

The auditory system of humanoid robots has gained increased attention in recent years. This system typically acquires the surrounding sound field by means of a microphone array. Signals acquired by the array are then processed using various…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-05 Vladimir Tourbabin , Boaz Rafaely

Robotic vision plays a major role in factory automation to service robot applications. However, the traditional use of frame-based camera sets a limitation on continuous visual feedback due to their low sampling rate and redundant data in…