中文
相关论文

相关论文: Multi-scale Multi-instance Visual Sound Localizati…

200 篇论文

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

声音 · 计算机科学 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

We propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound source localization…

声音 · 计算机科学 2018-09-18 Weipeng He , Petr Motlicek , Jean-Marc Odobez

Conventional approaches to sound localization and separation are based on microphone arrays in artificial systems. Inspired by the selective perception of human auditory system, we design a multi-source listening system which can separate…

声音 · 计算机科学 2019-11-11 Xuecong Sun , Han Jia , Zhe Zhang , Yuzhen Yang , Zhaoyong Sun , Jun Yang

Different self-supervised tasks (SSL) reveal different features from the data. The learned feature representations can exhibit different performance for each downstream task. In this light, this work aims to combine Multiple SSL tasks…

计算机视觉与模式识别 · 计算机科学 2022-01-05 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

The goal of visual answering localization (VAL) in the video is to obtain a relevant and concise time clip from a video as the answer to the given natural language question. Early methods are based on the interaction modelling between video…

计算机视觉与模式识别 · 计算机科学 2022-10-31 Yixuan Weng , Bin Li

Video segmentation -- partitioning video frames into multiple segments or objects -- plays a critical role in a broad range of practical applications, from enhancing visual effects in movie, to understanding scenes in autonomous driving, to…

计算机视觉与模式识别 · 计算机科学 2022-11-30 Tianfei Zhou , Fatih Porikli , David Crandall , Luc Van Gool , Wenguan Wang

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual…

We present a novel framework, Localized Image Stylization with Audio (LISA) which performs audio-driven localized image stylization. Sound often provides information about the specific context of the scene and is closely related to a…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Seung Hyun Lee , Chanyoung Kim , Wonmin Byeon , Sang Ho Yoon , Jinkyu Kim , Sangpil Kim

Multi-view subspace learning (MSL) aims to find a low-dimensional subspace of the data obtained from multiple views. Different from single view case, MSL should take both common and specific knowledge among different views into…

机器学习 · 计算机科学 2018-11-08 Hongwei Yong , Deyu Meng , Jinxing Li , Wangmeng Zuo , Lei Zhang

A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial…

The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them temporally through…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Syed Hesham Syed Ariff , Yun Liu , Guolei Sun , Jing Yang , Henghui Ding , Xue Geng , Xudong Jiang

Recognizing multiple objects in an image is challenging due to occlusions, and becomes even more so when the objects are small. While promising, existing multi-label image recognition models do not explicitly learn context-based…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Hasib Zunair , A. Ben Hamza

Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acousto-optic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction. Moreover, it is…

信号处理 · 电气工程与系统科学 2024-11-13 Risako Tanigawa , Kenji Ishikawa , Noboru Harada , Yasuhiro Oikawa

Visual place recognition (VPR) remains challenging due to significant viewpoint changes and appearance variations. Mainstream works tackle these challenges by developing various feature aggregation methods to transform deep features into…

计算机视觉与模式识别 · 计算机科学 2024-07-10 Teng Wang , Lingquan Meng , Lei Cheng , Changyin Sun

Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic computational cost, limiting their suitability for resource…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Qing Zhong , Guodong Ding , Lingqiao Liu , Zaiwen Feng , Lin Yuanbo Wu , Angela Yao

Large-scale pre-trained image-text models exhibit robust multimodal representations, yet applying the Contrastive Language-Image Pre-training (CLIP) model to audio-visual localization remains challenging. Replacing the classification token…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Khanh Binh Nguyen , Chae Jung Park

This paper proposes an attributable visual similarity learning (AVSL) framework for a more accurate and explainable similarity measure between images. Most existing similarity learning methods exacerbate the unexplainability by mapping each…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Borui Zhang , Wenzhao Zheng , Jie Zhou , Jiwen Lu

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, and to associate them with a visual signal. Our method jointly…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Xixi Hu , Ziyang Chen , Andrew Owens

Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling negatives in prior…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Zengjie Song , Jiangshe Zhang , Yuxi Wang , Junsong Fan , Zhaoxiang Zhang