中文
相关论文

相关论文: SoundSpaces 2.0: A Simulation Platform for Visual-…

200 篇论文

Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achieving autonomous navigation. The core challenge of this task…

声音 · 计算机科学 2026-04-06 Xinyu Zhou , Yinfeng Yu

We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images, and…

In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment…

声音 · 计算机科学 2025-10-24 Junjie Zheng , Gongyu Chen , Chaofan Ding , Zihao Chen

Sound plays a major role in human perception. Along with vision, it provides essential information for understanding our surroundings. Despite advances in neural implicit representations, learning acoustics that align with visual scenes…

声音 · 计算机科学 2025-10-03 Amandine Brunetto , Sascha Hornauer , Fabien Moutarde

Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However,…

机器人学 · 计算机科学 2024-03-04 Amandine Brunetto , Sascha Hornauer , Stella X. Yu , Fabien Moutarde

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural rendering approach:…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Changan Chen , Alexander Richard , Roman Shapovalov , Vamsi Krishna Ithapu , Natalia Neverova , Kristen Grauman , Andrea Vedaldi

Computational ultrasound imaging has become a well-established methodology in the ultrasound community. In the accompanying paper (part I), we described a new ultrasound simulator (SIMUS) for Matlab, which belongs to the Matlab UltraSound…

医学物理 · 物理学 2022-04-06 Amanda Cigier , François Varray , Damien Garcia

Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce a strategy to map…

音频与语音处理 · 电气工程与系统科学 2024-03-05 Xinmeng Xu , Yuhong Yang , Weiping Tu

Audio-visual navigation task requires an agent to find a sound source in a realistic, unmapped 3D environment by utilizing egocentric audio-visual observations. Existing audio-visual navigation works assume a clean environment that solely…

声音 · 计算机科学 2022-02-23 Yinfeng Yu , Wenbing Huang , Fuchun Sun , Changan Chen , Yikai Wang , Xiaohong Liu

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Bohan Li , Yasheng Sun , Jingxin Dong , Zheng Zhu , Jinming Liu , Xin Jin , Wenjun Zeng

Noise pollution investigation takes advantage of two common methods of diagnosis: measurement using a Sound Level Meter and acoustical imaging. The former enables a detailed analysis of the surrounding noise spectrum whereas the latter is…

We focus on the task of soundscape mapping, which involves predicting the most probable sounds that could be perceived at a particular geographic location. We utilise recent state-of-the-art models to encode geotagged audio, a textual…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Subash Khanal , Srikumar Sastry , Aayush Dhakal , Nathan Jacobs

Audio foundation models learn general-purpose audio representations that facilitate a wide range of downstream tasks. While the performance of these models has greatly increased for conventional single-channel, dry audio clips, their…

声音 · 计算机科学 2026-02-05 Goksenin Yuksel , Marcel van Gerven , Kiki van der Heijden

In today's tech-driven world, significant advancements in artificial intelligence and virtual reality have emerged. These developments drive research into exploring their intersection in the realm of soundscape. Not only do these…

声音 · 计算机科学 2025-04-11 Rima Ayoubi , Laurent Lescop , Sang Bum Park

Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acousto-optic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction. Moreover, it is…

信号处理 · 电气工程与系统科学 2024-11-13 Risako Tanigawa , Kenji Ishikawa , Noboru Harada , Yasuhiro Oikawa

We present the Geometric-Wave Acoustic (GWA) dataset, a large-scale audio dataset of about 2 million synthetic room impulse responses (IRs) and their corresponding detailed geometric and simulation configurations. Our dataset samples…

声音 · 计算机科学 2022-06-22 Zhenyu Tang , Rohith Aralikatti , Anton Ratnarajah , Dinesh Manocha

Recent neural audio codecs have achieved impressive reconstruction quality, typically relying on quantization methods such as Residual Vector Quantization (RVQ), Vector Quantization (VQ) and Finite Scalar Quantization (FSQ). However, these…

声音 · 计算机科学 2026-05-19 Tal Shuster , Eliya Nachmani

The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference.…

音频与语音处理 · 电气工程与系统科学 2025-12-17 Sungnyun Kim

We present a novel geometric deep learning method to compute the acoustic scattering properties of geometric objects. Our learning algorithm uses a point cloud representation of objects to compute the scattering properties and integrates…

声音 · 计算机科学 2021-05-19 Hsien-Yu Meng , Zhenyu Tang , Dinesh Manocha

The simulation of two-dimensional (2D) wave propagation is an affordable computational task and its use can potentially improve time performance in vocal tracts' acoustic analysis. Several models have been designed that rely on 2D wave…

声音 · 计算机科学 2019-09-23 Debasish Ray Mohapatra , Victor Zappi , Sidney Fels
‹ 上一页 1 8 9 10 下一页 ›