English
Related papers

Related papers: Learning Tri-modal Embeddings for Zero-Shot Sounds…

200 papers

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Recent audio-to-image models have shown impressive performance in generating images of specific objects conditioned on their corresponding sounds. However, these models fail to reconstruct real-world landscapes conditioned on environmental…

This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive…

Sound · Computer Science 2024-01-12 Zhengcong Fei , Mingyuan Fan , Junshi Huang

Language grounding aims at linking the symbolic representation of language (e.g., words) into the rich perceptual knowledge of the outside world. The general approach is to embed both textual and visual information into a common space -the…

Computation and Language · Computer Science 2021-09-15 Hassan Shahmohammadi , Hendrik P. A. Lensch , R. Harald Baayen

Audio perception is a key to solving a variety of problems ranging from acoustic scene analysis, music meta-data extraction, recommendation, synthesis and analysis. It can potentially also augment computers in doing tasks that humans do…

Sound · Computer Science 2020-02-12 Prateek Verma , Kenneth Salisbury

Zero-shot learning enables models to generalise to unseen classes by leveraging semantic information, bridging the gap between training and testing sets with non-overlapping classes. While much research has focused on zero-shot learning in…

Sound · Computer Science 2025-07-03 Ysobel Sims , Alexandre Mendes , Stephan Chalup

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Sonar-based indoor mapping systems have been widely employed in robotics for several decades. While such systems are still the mainstream in underwater and pipe inspection settings, the vulnerability to noise reduced, over time, their…

Robotics · Computer Science 2024-09-19 Usama Saqib , Letizia Marchegiani , Jesper Rindom Jensen

Mobile robots operating in unknown urban environments encounter a wide range of complex terrains to which they must adapt their planned trajectory for safe and efficient navigation. Most existing approaches utilize supervised learning to…

Robotics · Computer Science 2021-11-05 Jannik Zürn , Wolfram Burgard , Abhinav Valada

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-08 Davide Berghi , Philip J. B. Jackson

While interacting in the world is a multi-sensory experience, many robots continue to predominantly rely on visual perception to map and navigate in their environments. In this work, we propose Audio-Visual-Language Maps (AVLMaps), a…

Robotics · Computer Science 2023-03-28 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely drops in the presence of background sound sources. Our…

Sound · Computer Science 2025-06-06 Emiliano Acevedo , Martín Rocamora , Magdalena Fuentes

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across multiple modalities…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Rishit Dagli , Shivesh Prakash , Robert Wu , Houman Khosravani

Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leaving the spatial information in multi-channel recordings…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-16 Chenxin Yu , Hao Ma , Xu Li , Xiao-Lei Zhang , Mingjie Shao , Chi Zhang , Xuelong Li

Finding sound effects or environmental sounds that match a creator's intended impression remains a largely manual process in multimedia production. This is especially relevant for comics and other visual media, where visually stylized…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Keisuke Imoto , Yamato Kojima , Takao Tsuchiya

Several recent publications have proposed methods for mapping images into continuous semantic embedding spaces. In some cases the embedding space is trained jointly with the image transformation. In other cases the semantic embedding space…

Machine Learning · Computer Science 2017-02-28 Mohammad Norouzi , Tomas Mikolov , Samy Bengio , Yoram Singer , Jonathon Shlens , Andrea Frome , Greg S. Corrado , Jeffrey Dean

We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-01 Jaeyeon Kim , Jaeyoon Jung , Jinjoo Lee , Sang Hoon Woo

Few-shot audio-visual acoustics modeling seeks to synthesize the room impulse response in arbitrary locations with few-shot observations. To sufficiently exploit the provided few-shot data for accurate acoustic modeling, we present a…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Diwei Huang , Kunyang Lin , Peihao Chen , Qing Du , Mingkui Tan

Image geolocalization is the challenging task of predicting the geographic coordinates of origin for a given photo. It is an unsolved problem relying on the ability to combine visual clues with general knowledge about the world to make…

Computer Vision and Pattern Recognition · Computer Science 2023-02-02 Lukas Haas , Silas Alberti , Michal Skreta

With the ever-increasing amount of data, the central challenge in multimodal learning involves limitations of labelled samples. For the task of classification, techniques such as meta-learning, zero-shot learning, and few-shot learning…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Nihar Bendre , Kevin Desai , Peyman Najafirad