English
Related papers

Related papers: Multimodal Deep Learning Method for Real-Time Spat…

200 papers

Artificial reverberation (AR) models play a central role in various audio applications. Therefore, estimating the AR model parameters (ARPs) of a reference reverberation is a crucial task. Although a few recent deep-learning-based…

Sound · Computer Science 2022-07-21 Sungho Lee , Hyeong-Seok Choi , Kyogu Lee

Despite impressive advances in recent multimodal large language models (MLLMs), state-of-the-art models such as from the GPT-4 suite still struggle with knowledge-intensive tasks. To address this, we consider Reverse Image Retrieval (RIR)…

Computation and Language · Computer Science 2024-05-30 Jialiang Xu , Michael Moor , Jure Leskovec

We propose a mesh-based neural network (MESH2IR) to generate acoustic impulse responses (IRs) for indoor 3D scenes represented using a mesh. The IRs are used to create a high-quality sound experience in interactive applications and audio…

Sound · Computer Science 2022-07-13 Anton Ratnarajah , Zhenyu Tang , Rohith Chandrashekar Aralikatti , Dinesh Manocha

Channel state information (CSI) plays a critical role in achieving the potential benefits of massive multiple input multiple output (MIMO) systems. In frequency division duplex (FDD) massive MIMO systems, the base station (BS) relies on…

Information Theory · Computer Science 2022-10-03 Zhengyang Hu , Guanzhang Liu , Qi Xie , Jiang Xue , Deyu Meng , Deniz Gunduz

We introduce HiFi-HARP, a large-scale dataset of 7th-order Higher-Order Ambisonic Room Impulse Responses (HOA-RIRs) consisting of more than 100,000 RIRs generated via a hybrid acoustic simulation in realistic indoor scenes. HiFi-HARP…

Sound · Computer Science 2025-10-27 Shivam Saini , Jürgen Peissig

Implicit neural representations (INRs) are a rapidly growing research field, which provides alternative ways to represent multimedia signals. Recent applications of INRs include image super-resolution, compression of high-dimensional…

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide…

Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-31 Shaoheng Xu , Chunyi Sun , Jihui Zhang , Amy Bastine , Prasanga N. Samarasinghe , Thushara D. Abhayapala , Hongdong Li

This paper presents a reverberation module for source-filter-based neural vocoders that improves the performance of reverberant effect modeling. This module uses the output waveform of neural vocoders as an input and produces a reverberant…

Sound · Computer Science 2020-05-18 Yang Ai , Xin Wang , Junichi Yamagishi , Zhen-Hua Ling

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (RIRs), and thus offer…

Sound · Computer Science 2026-05-04 Akira Takahashi , Ryosuke Sawata , Shusuke Takahashi , Yuki Mitsufuji

Room equalisation aims to increase the quality of loudspeaker reproduction in reverberant environments, compensating for colouration caused by imperfect room reflections and frequency dependant loudspeaker directivity. A common technique in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 James Brooks-Park , Martin Bo Møller , Jan Østergaard , Søren Bech , Steven van de Par

In real-world acoustic scenarios, there often are multiple sound sources present in a room. These sources are situated in various locations and produce sounds that reach the listener from multiple directions. The presence of multiple…

Sound · Computer Science 2023-05-26 Kyungyun Lee , Jeonghun Seo , Keunwoo Choi , Sangmoon Lee , Ben Sangbae Chon

Current end-to-end deep Reinforcement Learning (RL) approaches require jointly learning perception, decision-making and low-level control from very sparse reward signals and high-dimensional inputs, with little capability of incorporating…

Machine Learning · Computer Science 2019-10-10 Vibhavari Dasagi , Robert Lee , Serena Mou , Jake Bruce , Niko Sünderhauf , Jürgen Leitner

This paper presents the initial stages in the development of a deep learning classifier for generalised Resident Space Object (RSO) characterisation that combines high-fidelity simulated light curves with transfer learning to improve the…

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Realtime 4D reconstruction for dynamic scenes remains a crucial challenge for autonomous driving perception. Most existing methods rely on depth estimation through self-supervision or multi-modality sensor fusion. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xin Fei , Wenzhao Zheng , Yueqi Duan , Wei Zhan , Masayoshi Tomizuka , Kurt Keutzer , Jiwen Lu

Precise spatial modeling in the operating room (OR) is foundational to many clinical tasks, supporting intraoperative awareness, hazard avoidance, and surgical decision-making. While existing approaches leverage large-scale multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Peiqi He , Zhenhao Zhang , Yixiang Zhang , Xiongjun Zhao , Shaoliang Peng

The perceptual evaluation of spatial audio algorithms is an important step in the development of immersive audio applications, as it ensures that synthesized sound fields meet quality standards in terms of listening experience, spatial…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-04 Paolo Ostan , Francesca Del Gaudio , Federico Miotello , Mirco Pezzoli , Fabio Antonacci

How to improve the ability of scene representation is a key issue in vision-oriented decision-making applications, and current approaches usually learn task-relevant state representations within visual reinforcement learning to address this…

Artificial Intelligence · Computer Science 2024-10-24 Dayang Liang , Jinyang Lai , Yunlong Liu

Dynamic MRI reconstruction, one of inverse problems, has seen a surge by the use of deep learning techniques. Especially, the practical difficulty of obtaining ground truth data has led to the emergence of unsupervised learning approaches.…

Image and Video Processing · Electrical Eng. & Systems 2025-10-28 Dayoung Baik , Jaejun Yoo
‹ Prev 1 3 4 5 6 7 10 Next ›