English
Related papers

Related papers: Learning Tri-modal Embeddings for Zero-Shot Sounds…

200 papers

Worldwide visual geo-localization aims to determine the geographic location of an image anywhere on Earth using only its visual content. Despite recent progress, learning expressive representations of geographic space remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Angel Daruna , Nicholas Meegan , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

We propose a weakly supervised approach for creating maps using free-form textual descriptions. We refer to this work of creating textual maps as zero-shot mapping. Prior works have approached mapping tasks by developing models that predict…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Aayush Dhakal , Adeel Ahmad , Subash Khanal , Srikumar Sastry , Hannah Kerner , Nathan Jacobs

Accurately localizing 3D sound sources and estimating their semantic labels -- where the sources may not be visible, but are assumed to lie on the physical surface of objects in the scene -- have many real applications, including detecting…

Sound · Computer Science 2024-12-31 Yuhang He , Sangyun Shin , Anoop Cherian , Niki Trigoni , Andrew Markham

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-12 Jinzheng Zhao , Yong Xu , Haohe Liu , Davide Berghi , Xinyuan Qian , Qiuqiang Kong , Junqi Zhao , Mark D. Plumbley , Wenwu Wang

Most speech recognition tasks pertain to mapping words across two modalities: acoustic and orthographic. In this work, we suggest learning encoders that map variable-length, acoustic or phonetic, sequences that represent words into…

Machine Learning · Computer Science 2019-08-02 Mohamed El-Geish

In the past, the rapidly evolving field of sound classification greatly benefited from the application of methods from other domains. Today, we observe the trend to fuse domain-specific tasks and approaches together, which provides the…

Sound · Computer Science 2022-09-12 Andrey Guzhov , Federico Raue , Jörn Hees , Andreas Dengel

Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the…

Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent…

Sound · Computer Science 2025-07-11 Haokun Tian , Stefan Lattner , Charalampos Saitis

Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio…

Sound · Computer Science 2025-07-08 Ludovic Tuncay , Etienne Labbé , Emmanouil Benetos , Thomas Pellegrini

Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these models excel in various downstream tasks, including visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Madhukar Reddy Vongala , Saurabh Srivastava , Jana Košecká

In this paper, we tackle the task of musical stem retrieval. Given a musical mix, it consists in retrieving a stem that would fit with it, i.e., that would sound pleasant if played together. To do so, we introduce a new method based on…

Sound · Computer Science 2025-02-25 Alain Riou , Antonin Gagneré , Gaëtan Hadjeres , Stefan Lattner , Geoffroy Peeters

Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models exhibit limitations in processing spatial audio and perceiving spatial acoustic scenes. To…

Sound · Computer Science 2025-09-19 Jinbo Hu , Yin Cao , Ming Wu , Zhenbo Luo , Jun Yang

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Rishabh Garg , Ruohan Gao , Kristen Grauman

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and negative pairs based on…

Machine Learning · Computer Science 2025-03-11 Theodor Lundqvist , Ludvig Delvret

Integrating visual and linguistic information into a single multimodal representation is an unsolved problem with wide-reaching applications to both natural language processing and computer vision. In this paper, we present a simple method…

Machine Learning · Statistics 2017-03-28 Guillem Collell , Teddy Zhang , Marie-Francine Moens

Visual geo-localization for drones faces critical degradation under weather perturbations, \eg, rain and fog, where existing methods struggle with two inherent limitations: 1) Heavy reliance on limited weather categories that constrain…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Jiahao Wen , Hang Yu , Zhedong Zheng

Audio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Jie Hong , Zeeshan Hayder , Junlin Han , Pengfei Fang , Mehrtash Harandi , Lars Petersson

Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping,…

Robotics · Computer Science 2025-06-10 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

We study the problem of compositional zero-shot learning for object-attribute recognition. Prior works use visual features extracted with a backbone network, pre-trained for object classification and thus do not capture the subtly distinct…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Nirat Saini , Khoi Pham , Abhinav Shrivastava

Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint…

Sound · Computer Science 2025-10-17 Qixin Deng , Bryan Pardo , Thrasyvoulos N Pappas