English
Related papers

Related papers: Multi-scale Multi-instance Visual Sound Localizati…

200 papers

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

We propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound source localization…

Sound · Computer Science 2018-09-18 Weipeng He , Petr Motlicek , Jean-Marc Odobez

Conventional approaches to sound localization and separation are based on microphone arrays in artificial systems. Inspired by the selective perception of human auditory system, we design a multi-source listening system which can separate…

Sound · Computer Science 2019-11-11 Xuecong Sun , Han Jia , Zhe Zhang , Yuzhen Yang , Zhaoyong Sun , Jun Yang

Different self-supervised tasks (SSL) reveal different features from the data. The learned feature representations can exhibit different performance for each downstream task. In this light, this work aims to combine Multiple SSL tasks…

Computer Vision and Pattern Recognition · Computer Science 2022-01-05 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

The goal of visual answering localization (VAL) in the video is to obtain a relevant and concise time clip from a video as the answer to the given natural language question. Early methods are based on the interaction modelling between video…

Computer Vision and Pattern Recognition · Computer Science 2022-10-31 Yixuan Weng , Bin Li

Video segmentation -- partitioning video frames into multiple segments or objects -- plays a critical role in a broad range of practical applications, from enhancing visual effects in movie, to understanding scenes in autonomous driving, to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Tianfei Zhou , Fatih Porikli , David Crandall , Luc Van Gool , Wenguan Wang

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual…

Machine Learning · Computer Science 2025-08-14 Duc Cao-Dinh , Khai Le-Duc , Anh Dao , Bach Phan Tat , Chris Ngo , Duy M. H. Nguyen , Nguyen X. Khanh , Thanh Nguyen-Tang

We present a novel framework, Localized Image Stylization with Audio (LISA) which performs audio-driven localized image stylization. Sound often provides information about the specific context of the scene and is closely related to a…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Seung Hyun Lee , Chanyoung Kim , Wonmin Byeon , Sang Ho Yoon , Jinkyu Kim , Sangpil Kim

Multi-view subspace learning (MSL) aims to find a low-dimensional subspace of the data obtained from multiple views. Different from single view case, MSL should take both common and specific knowledge among different views into…

Machine Learning · Computer Science 2018-11-08 Hongwei Yong , Deyu Meng , Jinxing Li , Wangmeng Zuo , Lei Zhang

A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial…

The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them temporally through…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Syed Hesham Syed Ariff , Yun Liu , Guolei Sun , Jing Yang , Henghui Ding , Xue Geng , Xudong Jiang

Recognizing multiple objects in an image is challenging due to occlusions, and becomes even more so when the objects are small. While promising, existing multi-label image recognition models do not explicitly learn context-based…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Hasib Zunair , A. Ben Hamza

Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acousto-optic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction. Moreover, it is…

Signal Processing · Electrical Eng. & Systems 2024-11-13 Risako Tanigawa , Kenji Ishikawa , Noboru Harada , Yasuhiro Oikawa

Visual place recognition (VPR) remains challenging due to significant viewpoint changes and appearance variations. Mainstream works tackle these challenges by developing various feature aggregation methods to transform deep features into…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Teng Wang , Lingquan Meng , Lei Cheng , Changyin Sun

Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic computational cost, limiting their suitability for resource…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Qing Zhong , Guodong Ding , Lingqiao Liu , Zaiwen Feng , Lin Yuanbo Wu , Angela Yao

Large-scale pre-trained image-text models exhibit robust multimodal representations, yet applying the Contrastive Language-Image Pre-training (CLIP) model to audio-visual localization remains challenging. Replacing the classification token…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Khanh Binh Nguyen , Chae Jung Park

This paper proposes an attributable visual similarity learning (AVSL) framework for a more accurate and explainable similarity measure between images. Most existing similarity learning methods exacerbate the unexplainability by mapping each…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Borui Zhang , Wenzhao Zheng , Jie Zhou , Jiwen Lu

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, and to associate them with a visual signal. Our method jointly…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Xixi Hu , Ziyang Chen , Andrew Owens

Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling negatives in prior…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Zengjie Song , Jiangshe Zhang , Yuxi Wang , Junsong Fan , Zhaoxiang Zhang