English
Related papers

Related papers: Multi-scale Multi-instance Visual Sound Localizati…

200 papers

This work addresses the task of weakly-supervised object localization. The goal is to learn object localization using only image-level class labels, which are much easier to obtain compared to bounding box annotations. This task is…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 David Kim , Sinhae Cha , Byeongkeun Kang

We propose a novel method for instance label segmentation of dense 3D voxel grids. We target volumetric scene representations, which have been acquired with depth sensors or multi-view stereo methods and which have been processed with…

Computer Vision and Pattern Recognition · Computer Science 2019-11-04 Jean Lahoud , Bernard Ghanem , Marc Pollefeys , Martin R. Oswald

Video action recognition is a challenging but important task for understanding and discovering what the video does. However, acquiring annotations for a video is costly, and semi-supervised learning (SSL) has been studied to improve…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Seokun Kang , Taehwan Kim

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Yuxin Guo , Shijie Ma , Hu Su , Zhiqing Wang , Yuhao Zhao , Wei Zou , Siyang Sun , Yun Zheng

The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, their handling of long-range dependencies struggles due to…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Sitong Gong , Yunzhi Zhuge , Lu Zhang , Yifan Wang , Pingping Zhang , Lijun Wang , Huchuan Lu

Online construction of open-ended language scenes is crucial for robotic applications, where open-vocabulary interactive scene understanding is required. Recently, neural implicit representation has provided a promising direction for online…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Muer Tie , Julong Wei , Zhengjun Wang , Ke Wu , Shansuai Yuan , Kaizhao Zhang , Jie Jia , Jieru Zhao , Zhongxue Gan , Wenchao Ding

Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-visual LLMs, and to humans using their ears, eyes, or both…

Sound · Computer Science 2025-05-13 Xilin Jiang , Junkai Wu , Vishal Choudhari , Nima Mesgarani

In the task of audio-visual sound source separation, which leverages visual information for sound source separation, identifying objects in an image is a crucial step prior to separating the sound source. However, existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Takashi Oya , Shohei Iwase , Shigeo Morishima

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

We present a novel 3D instance segmentation framework for Multi-View Stereo (MVS) buildings in urban scenes. Unlike existing works focusing on semantic segmentation of urban scenes, the emphasis of this work lies in detecting and segmenting…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Jiazhou Chen , Yanghui Xu , Shufang Lu , Ronghua Liang , Liangliang Nan

Visual speech recognition (VSR) is the task of recognizing spoken language from video input only, without any audio. VSR has many applications as an assistive technology, especially if it could be deployed in mobile devices and embedded…

Computation and Language · Computer Science 2019-06-06 Nilay Shrivastava , Astitwa Saxena , Yaman Kumar , Rajiv Ratn Shah , Debanjan Mahata , Amanda Stent

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

Computer Vision and Pattern Recognition · Computer Science 2017-12-21 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

Locating specific segments within an instructional video is an efficient way to acquire guiding knowledge. Generally, the task of obtaining video segments for both verbal explanations and visual demonstrations is known as visual answer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Chang Zong , Bin Li , Shoujun Zhou , Jian Wan , Lei Zhang

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Recent open-vocabulary 3D scene understanding approaches mainly focus on training 3D networks through contrastive learning with point-text pairs or by distilling 2D features into 3D models via point-pixel alignment. While these methods show…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Xingyilang Yin , Jiale Wang , Xi Yang , Mutian Xu , Xu Gu , Nannan Wang

Open-vocabulary image semantic segmentation (OVS) seeks to segment images into semantic regions across an open set of categories. Existing OVS methods commonly depend on foundational vision-language models and utilize similarity computation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Qinglong Cao , Yuntian Chen , Chao Ma , Xiaokang Yang

Over the past few years, there has been a great deal of research on navigation tasks in indoor environments using deep reinforcement learning agents. Most of these tasks use only visual information in the form of first-person images to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Haru Kondoh , Asako Kanezaki

Visual localization remains challenging in dynamic environments where fluctuating lighting, adverse weather, and moving objects disrupt appearance cues. Despite advances in feature representation, current absolute pose regression methods…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Zhongtao Tian , Wenhao Huang , Zhidong Chen , Xiao Wei Sun

The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios are usually contaminated by off-screen sound and background…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Xian Liu , Rui Qian , Hang Zhou , Di Hu , Weiyao Lin , Ziwei Liu , Bolei Zhou , Xiaowei Zhou

A long-standing goal in the field of sensory substitution is to enable sound perception for deaf and hard of hearing (DHH) people by visualizing audio content. Different from existing models that translate to hand sign language, between…

Human-Computer Interaction · Computer Science 2023-02-15 Chunjin Song , Yuchi Zhang , Willis Peng , Parmis Mohaghegh , Bastian Wandt , Helge Rhodin