中文
相关论文

相关论文: MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics …

200 篇论文

Few-shot segmentation aims to segment unseen object categories from just a handful of annotated examples. This requires mechanisms that can both identify semantically related objects across images and accurately produce segmentation masks.…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Claudia Cuttano , Gabriele Trivigno , Giuseppe Averta , Carlo Masone

Detecting the presence of animal vocalisations in nature is essential to study animal populations and their behaviors. A recent development in the field is the introduction of the task known as few-shot bioacoustic sound event detection,…

音频与语音处理 · 电气工程与系统科学 2024-03-28 Jinhua Liang , Ines Nolasco , Burooj Ghani , Huy Phan , Emmanouil Benetos , Dan Stowell

In this paper, we introduce a novel attentional similarity module for the problem of few-shot sound recognition. Given a few examples of an unseen sound event, a classifier must be quickly adapted to recognize the new sound event without…

声音 · 计算机科学 2019-02-19 Szu-Yu Chou , Kai-Hsiang Cheng , Jyh-Shing Roger Jang , Yi-Hsuan Yang

In this paper our objectives are, first, networks that can embed audio and visual inputs into a common space that is suitable for cross-modal retrieval; and second, a network that can localize the object that sounds in an image, given the…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Relja Arandjelović , Andrew Zisserman

Popular approaches for few-shot classification consist of first learning a generic data representation based on a large annotated dataset, before adapting the representation to new classes given only a few labeled samples. In this work, we…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Nikita Dvornik , Cordelia Schmid , Julien Mairal

Recognition of remote sensing (RS) or aerial images is currently of great interest, and advancements in deep learning algorithms added flavor to it in recent years. Occlusion, intra-class variance, lighting, etc., might arise while training…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Ankit Jha , Debabrata Pal , Mainak Singha , Naman Agarwal , Biplab Banerjee

We consider the problem of few-shot scene adaptive crowd counting. Given a target camera scene, our goal is to adapt a model to this specific scene with only a few labeled images of that scene. The solution to this problem has potential…

计算机视觉与模式识别 · 计算机科学 2020-06-22 Mahesh Kumar Krishna Reddy , Mohammad Hossain , Mrigank Rochan , Yang Wang

Semantic mapping is the incremental process of "mapping" relevant information of the world (i.e., spatial information, temporal events, agents and actions) to a formal description supported by a reasoning engine. Current research focuses on…

机器人学 · 计算机科学 2016-06-14 Roberto Capobianco , Jacopo Serafin , Johann Dichtl , Giorgio Grisetti , Luca Iocchi , Daniele Nardi

This paper proposes a real-time system integrating an acoustic material estimation from visual appearance and an on-the-fly mapping in the 3-dimension. The proposed method estimates the acoustic materials of surroundings in indoor scenes…

机器人学 · 计算机科学 2019-09-17 Taeyoung Kim , Youngsun Kwon , Sung-eui Yoon

To autonomously navigate in real-world environments, special in search and rescue operations, Unmanned Aerial Vehicles (UAVs) necessitate comprehensive maps to ensure safety. However, the prevalent metric map often lacks semantic…

机器人学 · 计算机科学 2024-01-17 Thanh Nguyen Canh , Armagan Elibol , Nak Young Chong , Xiem HoangVan

In recent years, zero-shot and few-shot learning in visual grounding have garnered considerable attention, largely due to the success of large-scale vision-language pre-training on expansive datasets such as LAION-5B and DataComp-1B.…

人工智能 · 计算机科学 2024-10-07 Sen Jia , Lei Li

For augmented (AR) and virtual reality (VR) applications, accurate estimates of the acoustic characteristics of a scene are critical for creating a sense of immersion. However, directly estimating Room-impulse Responses (RIRs) from scene…

音频与语音处理 · 电气工程与系统科学 2025-11-20 Ricardo Falcon-Perez , Ruohan Gao , Gregor Mueckl , Sebastia V. Amengual Gari , Ishwarya Ananthabhotla

Few-shot semantic segmentation aims to segment objects from previously unseen classes using only a limited number of labeled examples. In this paper, we introduce Label Anything, a novel transformer-based architecture designed for…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Pasquale De Marinis , Nicola Fanelli , Raffaele Scaringi , Emanuele Colonna , Giuseppe Fiameni , Gennaro Vessio , Giovanna Castellano

We study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition…

机器学习 · 计算机科学 2020-02-24 Bowen Shi , Ming Sun , Krishna C. Puvvada , Chieh-Chi Kao , Spyros Matsoukas , Chao Wang

Deep learning-based approaches to musical source separation are often limited to the instrument classes that the models are trained on and do not generalize to separate unseen instruments. To address this, we propose a few-shot musical…

声音 · 计算机科学 2022-05-04 Yu Wang , Daniel Stoller , Rachel M. Bittner , Juan Pablo Bello

We propose a few-shot learning method for spatial regression. Although Gaussian processes (GPs) have been successfully used for spatial regression, they require many observations in the target task to achieve a high predictive performance.…

机器学习 · 统计学 2020-10-12 Tomoharu Iwata , Yusuke Tanaka

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing,…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Zilong Deng , Federico Tombari , Marc Pollefeys , Johanna Wald , Daniel Barath

Speech-based machine learning (ML) has been heralded as a promising solution for tracking prosodic and spectrotemporal patterns in real-life that are indicative of emotional changes, providing a valuable window into one's cognitive and…

机器学习 · 计算机科学 2021-09-08 Kexin Feng , Theodora Chaspari

This paper introduces a new paradigm for sound source lo-calization referred to as virtual acoustic space traveling (VAST) and presents a first dataset designed for this purpose. Existing sound source localization methods are either based…

声音 · 计算机科学 2016-12-20 Clément Gaultier , Saurabh Kataria , Antoine Deleforge

We present a novel method for few-shot video classification, which performs appearance and temporal alignments. In particular, given a pair of query and support videos, we conduct appearance alignment via frame-level feature matching to…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Khoi D. Nguyen , Quoc-Huy Tran , Khoi Nguyen , Binh-Son Hua , Rang Nguyen