English
Related papers

Related papers: SoundSculpt: Direction and Semantics Driven Ambiso…

200 papers

Ambisonics is a scene-based spatial audio format that has several useful features compared to object-based formats, such as efficient whole scene rotation and versatility. However, it does not provide direct access to the individual source…

Sound · Computer Science 2023-06-21 Francesc Lluís , Nils Meyer-Kahlen , Vasileios Chatziioannou , Alex Hofmann

Supervised learning methods can solve the given problem in the presence of a large set of labeled data. However, the acquisition of a dataset covering all the target classes typically requires manual labeling which is expensive and…

Sound · Computer Science 2022-06-13 Duygu Dogan , Huang Xie , Toni Heittola , Tuomas Virtanen

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Sound field decomposition predicts waveforms in arbitrary directions using signals from a limited number of microphones as inputs. Sound field decomposition is fundamental to downstream tasks, including source localization, source…

Sound · Computer Science 2022-10-25 Qiuqiang Kong , Shilei Liu , Junjie Shi , Xuzhou Ye , Yin Cao , Qiaoxi Zhu , Yong Xu , Yuxuan Wang

Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and car honks. We…

Sound · Computer Science 2023-11-02 Bandhav Veluri , Malek Itani , Justin Chan , Takuya Yoshioka , Shyamnath Gollakota

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models exhibit limitations in processing spatial audio and perceiving spatial acoustic scenes. To…

Sound · Computer Science 2025-09-19 Jinbo Hu , Yin Cao , Ming Wu , Zhenbo Luo , Jun Yang

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environment make sounds…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Changan Chen , Ziad Al-Halah , Kristen Grauman

We propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speaker in a mixture. Typical approaches have been exploiting…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-26 Yasunori Ohishi , Marc Delcroix , Tsubasa Ochiai , Shoko Araki , Daiki Takeuchi , Daisuke Niizumi , Akisato Kimura , Noboru Harada , Kunio Kashino

Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application of these models,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Sooyoung Park , Arda Senocak , Joon Son Chung

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Yue Qiao , Vinay Kothapally , Meng Yu , Dong Yu

Ambisonics i.e., a full-sphere surround sound, is quintessential with 360-degree visual content to provide a realistic virtual reality (VR) experience. While 360-degree visual content capture gained a tremendous boost recently, the…

Sound · Computer Science 2019-08-20 Aakanksha Rana , Cagri Ozcinar , Aljoscha Smolic

In this paper our objectives are, first, networks that can embed audio and visual inputs into a common space that is suitable for cross-modal retrieval; and second, a network that can localize the object that sounds in an image, given the…

Computer Vision and Pattern Recognition · Computer Science 2018-07-27 Relja Arandjelović , Andrew Zisserman

Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leaving the spatial information in multi-channel recordings…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-16 Chenxin Yu , Hao Ma , Xu Li , Xiao-Lei Zhang , Mingjie Shao , Chi Zhang , Xuelong Li

Contrastive language--audio pretraining (CLAP) has achieved remarkable success as an audio--text embedding framework, but existing approaches are limited to monaural or single-source conditions and cannot fully capture spatial information.…

Ambisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-14 Ali Fallah , Shun Nakamura , Steven van de Par

Multichannel speech enhancement leverages spatial cues to improve intelligibility and quality, but most learning-based methods rely on specific microphone array geometry, unable to account for geometry changes. To mitigate this limitation,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Michael Tatarjitzky , Boaz Rafaely

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy…

Graphics · Computer Science 2021-12-02 Seung Hyun Lee , Wonseok Roh , Wonmin Byeon , Sang Ho Yoon , Chan Young Kim , Jinkyu Kim , Sangpil Kim

The importance of the information in the direct sound to human perception of spatial sound sources is an ongoing research topic. The classification between direct sound and diffuse or reverberant sound forms the basis of numerous studies in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-03 Eran Miller , Boaz Rafaely

Recently, deep learning-based beamforming algorithms have shown promising performance in target speech extraction tasks. However, most systems do not fully utilize spatial information. In this paper, we propose a target speech extraction…

Sound · Computer Science 2023-06-29 Aoqi Guo , Junnan Wu , Peng Gao , Wenbo Zhu , Qinwen Guo , Dazhi Gao , Yujun Wang
‹ Prev 1 2 3 10 Next ›