English
Related papers

Related papers: Scene-Aware Audio for 360\textdegree{} Videos

200 papers

Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or…

Sound · Computer Science 2025-07-14 Cheng Chi , Xiaoyu Li , Yuxuan Ke , Qunping Ni , Yao Ge , Xiaodong Li , Chengshi Zheng

Accurate intrinsic and extrinsic camera calibration can be an important prerequisite for robotic applications that rely on vision as input. While there is ongoing research on enabling camera calibration using natural images, many systems in…

Robotics · Computer Science 2025-04-16 Timm Linder , Kadir Yilmaz , David B. Adrian , Bastian Leibe

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning…

The topic of room equalisation has been at the forefront of research and product development for many years, with the aim of increasing the playback quality of loudspeakers in reverberant rooms. Traditional room equalisation systems…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 James Brooks-Park , Steven van de Par

We present a novel approach that improves the performance of reverberant speech separation. Our approach is based on an accurate geometric acoustic simulator (GAS) which generates realistic room impulse responses (RIRs) by modeling both…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-21 Rohith Aralikatti , Anton Ratnarajah , Zhenyu Tang , Dinesh Manocha

In multi-room environments, modelling the sound propagation is complex due to the coupling of rooms and diverse source-receiver positions. A common scenario is when the source and the receiver are in different rooms without a clear line of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-19 Kyung Yun Lee , Nils Meyer-Kahlen , Georg Götz , U. Peter Svensson , Sebastian J. Schlecht , Vesa Välimäki

Speech separation approaches for single-channel, dry speech mixtures have significantly improved. However, real-world spatial and reverberant acoustic environments remain challenging, limiting the effectiveness of these approaches for…

Room acoustic synthesis can be used in Virtual Reality (VR), Augmented Reality (AR) and gaming applications to enhance listeners' sense of immersion, realism and externalisation. A common approach is to use Geometrical Acoustics (GA) models…

Sound · Computer Science 2024-07-30 Matteo Scerbo , Lauri Savioja , Enzo De Sena

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Yuzhu Wang , Archontis Politis , Tuomas Virtanen

We propose a straightforward and cost-effective method to perform diffuse soundfield measurements for calibrating the magnitude response of a microphone array. Typically, such calibration is performed in a diffuse soundfield created in…

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate…

Computer Vision and Pattern Recognition · Computer Science 2018-02-13 Aviv Gabbay , Ariel Ephrat , Tavi Halperin , Shmuel Peleg

Robust spatial audio control relies on accurate acoustic propagation models, yet environmental variations, especially changes in the speed of sound, cause systematic mismatches that degrade performance. Existing methods either assume known…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-13 Andreas Jonas Fuglsig , Mads Græsbøll Christensen , Jesper Rindom Jensen

With the development of autonomous driving technology, sensor calibration has become a key technology to achieve accurate perception fusion and localization. Accurate calibration of the sensors ensures that each sensor can function properly…

Robotics · Computer Science 2023-05-29 Jixiang Li , Jiahao Pi , Guohang Yan , Yikang Li

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

This paper presents a complete hardware and software pipeline for real-time speech enhancement in noisy and reverberant conditions. The device consists of a microphone array and a camera mounted on eyeglasses, connected to an embedded…

Geometrical approaches for room acoustics simulation have the advantage of requiring limited computational resources while still achieving a high perceptual plausibility. A common approach is using the image source model for direct and…

Sound · Computer Science 2024-10-28 Siegfried Gündert , Stephan D. Ewert , Steven van de Par

In daily life, social interaction and acoustic communication often take place in complex acoustic environments (CAE) with a variety of interfering sounds and reverberation. For hearing research and the evaluation of hearing systems,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-30 Stefan Fichna , Thomas Biberger , Bernhard U. Seeber , Stephan D. Ewert

Room acoustics measurements are used in many areas of audio research, from physical acoustics modelling and speech enhancement to virtual reality applications. This paper documents the technical specifications and choices made in the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-24 Thomas McKenzie , Leo McCormack , Christoph Hold

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

Computer Vision and Pattern Recognition · Computer Science 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Automatic tuning of reverberation algorithms relies on the optimization of a cost function. While general audio similarity metrics are useful, they are not optimized for the specific statistical properties of reverberation in rooms. This…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-19 Gloria Dal Santo , Karolina Prawda , Sebastian J. Schlecht , Vesa Välimäki