中文
相关论文

相关论文: Structure from Silence: Learning Scene Structure f…

200 篇论文

For augmented (AR) and virtual reality (VR) applications, accurate estimates of the acoustic characteristics of a scene are critical for creating a sense of immersion. However, directly estimating Room-impulse Responses (RIRs) from scene…

音频与语音处理 · 电气工程与系统科学 2025-11-20 Ricardo Falcon-Perez , Ruohan Gao , Gregor Mueckl , Sebastia V. Amengual Gari , Ishwarya Ananthabhotla

There exists an unequivocal distinction between the sound produced by a static source and that produced by a moving one, especially when the source moves towards or away from the microphone. In this paper, we propose to use this connection…

声音 · 计算机科学 2022-11-01 Moitreya Chatterjee , Narendra Ahuja , Anoop Cherian

Echo-location is a broad approach to imaging and sensing that includes both man-made RADAR, LIDAR, SONAR and also animal navigation. However, full 3D information based on echo-location requires some form of scanning of the scene in order to…

图像与视频处理 · 电气工程与系统科学 2021-06-16 Alex Turpin , Valentin Kapitany , Jack Radford , Davide Rovelli , Kevin Mitchell , Ashley Lyons , Ilya Starshynov , Daniele Faccio

This work introduces a neural architecture for learning forward models of stochastic environments. The task is achieved solely through learning from temporal unstructured observations in the form of images. Once trained, the model allows…

机器学习 · 计算机科学 2021-12-16 Marian Andrecki , Nicholas K. Taylor

Accurately localizing 3D sound sources and estimating their semantic labels -- where the sources may not be visible, but are assumed to lie on the physical surface of objects in the scene -- have many real applications, including detecting…

声音 · 计算机科学 2024-12-31 Yuhang He , Sangyun Shin , Anoop Cherian , Niki Trigoni , Andrew Markham

We present an indoor acoustic simulation framework that supports both ultrasonic and audible signaling. The framework opens the opportunity for fast indoor acoustic data generation and positioning development. The improved…

音频与语音处理 · 电气工程与系统科学 2023-06-22 Daan Delabie , Chesney Buyle , Bert Cox , Liesbet Van der Perre , Lieven De Strycker

Autonomous soundscape augmentation systems typically use trained models to pick optimal maskers to effect a desired perceptual change. While acoustic information is paramount to such systems, contextual information, including participant…

声音 · 计算机科学 2024-07-03 Kenneth Ooi , Karn N. Watcharasupat , Bhan Lam , Zhen-Ting Ong , Woon-Seng Gan

Indoor scene recognition is a growing field with great potential for behaviour understanding, robot localization, and elderly monitoring, among others. In this study, we approach the task of scene recognition from a novel standpoint, using…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Andreea Glavan , Estefania Talavera

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

声音 · 计算机科学 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

The motivation of our research is to explore the possibilities of automatic sound-to-image (S2I) translation for enabling a human receiver to visually infer the occurrence of sound related events. We expect the computer to 'imagine' the…

声音 · 计算机科学 2022-03-10 Leonardo A. Fanzeres , Climent Nadeu

Audio-visual navigation combines sight and hearing to navigate to a sound-emitting source in an unmapped environment. While recent approaches have demonstrated the benefits of audio input to detect and find the goal, they focus on clean and…

声音 · 计算机科学 2023-01-04 Abdelrahman Younes , Daniel Honerkamp , Tim Welschehold , Abhinav Valada

Environmental soundscapes convey substantial ecological and social information regarding urban environments; however, their potential remains largely untapped in large-scale geographic analysis. In this study, we investigate the extent to…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Pengyu Chen , Xiao Huang , Teng Fei , Sicheng Wang

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Audio scene cartography for real or simulated stereo recordings is presented. This audio scene analysis is performed doing successively: a perceptive 10-subbands analysis, calculation of temporal laws for relative delays and gains between…

音频与语音处理 · 电气工程与系统科学 2024-01-10 Laurent Millot , Gérard Pelé , Mohammed Elliq

The localization of sound sources by the human brain is computationally simulated from a neurobiological perspective. The simulation includes the neural representation of temporal differences in acoustic signals between the ipsilateral and…

神经元与认知 · 定量生物学 2008-10-31 Nikesh S. Dattani

For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR…

声音 · 计算机科学 2024-09-24 Francesc Lluís , Nils Meyer-Kahlen

When we hear the word "house", we don't just process sound, we imagine walls, doors, memories. The brain builds meaning through layers, moving from raw acoustics to rich, multimodal associations. Inspired by this, we build on recent work…

机器学习 · 计算机科学 2025-11-11 Kateryna Shapovalenko , Quentin Auster

Analysis of respiratory sounds increases its importance every day. Many different methods are available in the analysis, and new techniques are continuing to be developed to further improve these methods. Features are extracted from audio…

声音 · 计算机科学 2021-01-22 Osman Balli , Yakup Kutlu

Knowing the geometry of a space is desirable for many applications, e.g. sound source localization, sound field reproduction or auralization. In circumstances where only acoustic signals can be obtained, estimating the geometry of a room is…

声音 · 计算机科学 2019-07-03 Linh Nguyen , Jaime Valls Miro , Xiaojun Qiu

World models have demonstrated impressive performance on robotic learning tasks. Many such tasks inherently demand multimodal reasoning; for example, filling a bottle with water will lead to visual information alone being ambiguous or…

机器人学 · 计算机科学 2025-12-10 Fan Zhang , Michael Gienger