中文
相关论文

相关论文: MRSAudio: A Large-Scale Multimodal Recorded Spatia…

200 篇论文

This paper introduces a new paradigm for sound source lo-calization referred to as virtual acoustic space traveling (VAST) and presents a first dataset designed for this purpose. Existing sound source localization methods are either based…

声音 · 计算机科学 2016-12-20 Clément Gaultier , Saurabh Kataria , Antoine Deleforge

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the…

计算机视觉与模式识别 · 计算机科学 2020-06-15 Karren Yang , Bryan Russell , Justin Salamon

Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural audio, which mimics human hearing, can capture richer spatial…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Zexu Pan , Shengkui Zhao , Yukun Ma , Haoxu Wang , Yiheng Jiang , Biao Tian , Bin Ma

Spatial Reasoning from language is essential for natural language understanding. Supporting it requires a representation scheme that can capture spatial phenomena encountered in language as well as in images and videos. Existing spatial…

计算与语言 · 计算机科学 2020-07-21 Soham Dan , Parisa Kordjamshidi , Julia Bonn , Archna Bhatia , Jon Cai , Martha Palmer , Dan Roth

This paper proposes a new task called spatial voice conversion, which aims to convert a target voice while preserving spatial information and non-target signals. Traditional voice conversion methods focus on single-channel waveforms,…

Augmented reality (AR) requires the seamless integration of visual, auditory, and linguistic channels for optimized human-computer interaction. While auditory and visual inputs facilitate real-time and contextual user guidance, the…

计算与语言 · 计算机科学 2023-10-19 Jing Bi , Nguyen Manh Nguyen , Ali Vosoughi , Chenliang Xu

Audio-visual speech enhancement (AVSE) has been found to be particularly useful at low signal-to-noise (SNR) ratios due to the immunity of the visual features to acoustic noise. However, a significant gap exists in AVSE methods tailored to…

音频与语音处理 · 电气工程与系统科学 2025-10-21 Danielle Yaffe , Ferdinand Campe , Prachi Sharma , Dorothea Kolossa , Boaz Rafaely

Binaural stereo audio is recorded by imitating the way the human ear receives sound, which provides people with an immersive listening experience. Existing approaches leverage autoencoders and directly exploit visual spatial information to…

声音 · 计算机科学 2023-11-15 Zhaojian Li , Bin Zhao , Yuan Yuan

Immersive environments have gradually become standard for visualizing and analyzing large or complex datasets that would otherwise be cumbersome, if not impossible, to explore through smaller scale computing devices. However, this type of…

人机交互 · 计算机科学 2019-03-12 Marco Cavallo , Mishal Dholakia , Matous Havlena , Kenneth Ocheltree , Mark Podlaseck

In this paper, we introduce a novel framework for spatial audio understanding of first-order ambisonic (FOA) signals through a question answering (QA) paradigm, aiming to extend the scope of sound event localization and detection (SELD)…

声音 · 计算机科学 2025-07-15 Parthasaarathy Sudarsanam , Archontis Politis

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Xikun Lu , Yunda Chen , Zehua Chen , Jie Wang , Mingxing Liu , Hongmei Hu , Chengshi Zheng , Stefan Bleeck , Jinqiu Sang

The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Jiaxin Huang , Ziwen Li , Hanlve Zhang , Runnan Chen , Xiao He , Yandong Guo , Wenping Wang , Tongliang Liu , Mingming Gong

In their everyday life, the speech recognition performance of human listeners is influenced by diverse factors, such as the acoustic environment, the talker and listener positions, possibly impaired hearing, and optional hearing devices.…

音频与语音处理 · 电气工程与系统科学 2021-04-02 Marc René Schädler

Data availability is essential in the development of acoustic signal processing algorithms, especially when it comes to data-driven approaches that demand large and diverse training datasets. For this reason, an increasing number of…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Stefano Damiano , Kathleen MacWilliam , Valerio Lorenzoni , Thomas Dietzen , Toon van Waterschoot

State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While…

人工智能 · 计算机科学 2025-10-17 Supriti Sinhamahapatra , Jan Niehues

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However,…

Environmental soundscapes convey substantial ecological and social information regarding urban environments; however, their potential remains largely untapped in large-scale geographic analysis. In this study, we investigate the extent to…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Pengyu Chen , Xiao Huang , Teng Fei , Sicheng Wang

Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping,…

机器人学 · 计算机科学 2025-06-10 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

We introduce a dataset for facilitating audio-visual analysis of music performances. The dataset comprises 44 simple multi-instrument classical music pieces assembled from coordinated but separately recorded performances of individual…

多媒体 · 计算机科学 2018-08-09 Bochen Li , Xinzhao Liu , Karthik Dinesh , Zhiyao Duan , Gaurav Sharma

Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However,…

机器人学 · 计算机科学 2024-03-04 Amandine Brunetto , Sascha Hornauer , Stella X. Yu , Fabien Moutarde