中文
相关论文

相关论文: PSM: Learning Probabilistic Embeddings for Multi-s…

200 篇论文

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed…

计算机视觉与模式识别 · 计算机科学 2021-02-12 Changan Chen , Sagnik Majumder , Ziad Al-Halah , Ruohan Gao , Santhosh Kumar Ramakrishnan , Kristen Grauman

From whirling ceiling fans to ticking clocks, the sounds that we hear subtly vary as we move through a scene. We ask whether these ambient sounds convey information about 3D scene structure and, if so, whether they provide a useful learning…

声音 · 计算机科学 2021-11-11 Ziyang Chen , Xixi Hu , Andrew Owens

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

计算机视觉与模式识别 · 计算机科学 2016-12-06 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine to have the same…

The task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understanding. Despite advancements in state-of-the-art (SOTA) models,…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Xavier Juanola , Gloria Haro , Magdalena Fuentes

Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies…

音频与语音处理 · 电气工程与系统科学 2024-11-04 Alan Pawlak , Hyunkook Lee , Aki Mäkivirta , Thomas Lund

We present a fast, scalable, and accurate Simultaneous Localization and Mapping (SLAM) system that represents indoor scenes as a graph of objects. Leveraging the observation that artificial environments are structured and occupied by…

机器人学 · 计算机科学 2020-11-06 Akash Sharma , Wei Dong , Michael Kaess

In this paper, we present a semantic mapping approach with multiple hypothesis tracking for data association. As semantic information has the potential to overcome ambiguity in measurements and place recognition, it forms an eminent…

机器人学 · 计算机科学 2020-12-09 Lukas Bernreiter , Abel Gawel , Hannes Sommer , Juan Nieto , Roland Siegwart , Cesar Cadena

One of the biggest challenges of acoustic scene classification (ASC) is to find proper features to better represent and characterize environmental sounds. Environmental sounds generally involve more sound sources while exhibiting less…

声音 · 计算机科学 2019-04-11 Hongwei Song , Jiqing Han , Shiwen Deng

In a sensor network with remote sensor devices, it is important to have a method that can accurately localize a sound event with a small amount of data transmitted from the sensors. In this paper, we propose a novel method for localization…

声音 · 计算机科学 2013-03-01 Hong Jiang , Boyd Mathews , Paul Wilford

Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationships among object parts. To address this limitation, we introduce Spatial Prediction (SP),…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Yang Shen , Yusen Cai , Weronika Hryniewska-Guzik , Qing Lin , Mengmi Zhang

In cluttered environments where visual sensors encounter heavy occlusion, such as in agricultural settings, tactile signals can provide crucial spatial information for the robot to locate rigid objects and maneuver around them. We introduce…

机器人学 · 计算机科学 2024-12-16 Moonyoung Lee , Uksang Yoo , Jean Oh , Jeffrey Ichnowski , George Kantor , Oliver Kroemer

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound…

人机交互 · 计算机科学 2026-01-27 Daehwa Kim , Chris Harrison

State-space models (SSMs) have emerged as a powerful foundation for long-range sequence modeling, with the HiPPO framework showing that continuous-time projection operators can be used to derive stable, memory-efficient dynamical systems…

机器学习 · 计算机科学 2026-02-27 Ruben Solozabal , Velibor Bojkovic , Hilal Alquabeh , Klea Ziu , Kentaro Inui , Martin Takac

We study the problem of localizing a configuration of points and planes from the collection of point-to-plane distances. This problem models simultaneous localization and mapping from acoustic echoes as well as the notable "structure from…

计算几何 · 计算机科学 2020-06-24 Miranda Krekovic , Ivan Dokmanic , Martin Vetterli

Having knowledge of the environmental context of the user i.e. the knowledge of the users' indoor location and the semantics of their environment, can facilitate the development of many of location-aware applications. In this paper, we…

声音 · 计算机科学 2018-04-03 Muhammad A. Shah , Bhiksha Raj , Khaled A. Harras

Our environment is filled with rich and dynamic acoustic information. When we walk into a cathedral, the reverberations as much as appearance inform us of the sanctuary's wide open space. Similarly, as an object moves around us, we expect…

声音 · 计算机科学 2023-01-18 Andrew Luo , Yilun Du , Michael J. Tarr , Joshua B. Tenenbaum , Antonio Torralba , Chuang Gan

Simultaneous Localization and Mapping (SLAM) techniques play a key role towards long-term autonomy of mobile robots due to the ability to correct localization errors and produce consistent maps of an environment over time. Contrarily to…

The objective of this work is to localize the sound sources in visual scenes. Existing audio-visual works employ contrastive learning by assigning corresponding audio-visual pairs from the same source as positives while randomly mismatched…

计算机视觉与模式识别 · 计算机科学 2022-02-08 Arda Senocak , Hyeonggon Ryu , Junsik Kim , In So Kweon
‹ 上一页 1 8 9 10 下一页 ›