中文
相关论文

相关论文: PSM: Learning Probabilistic Embeddings for Multi-s…

200 篇论文

The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Michael Bloesch , Jan Czarnowski , Ronald Clark , Stefan Leutenegger , Andrew J. Davison

Human-robot interaction requires a common understanding of the operational environment, which can be provided by a representation that blends geometric and symbolic knowledge: a semantic map. Through a semantic map the robot can interpret…

机器人学 · 计算机科学 2021-05-18 Sara Kaszuba , Sandeep Reddy Sabbella , Vincenzo Suriani , Francesco Riccio , Daniele Nardi

Deep learning-based sound event localization and classification is an emerging research area within wireless acoustic sensor networks. However, current methods for sound event localization and classification typically rely on a single…

Spatial-temporal forecasting plays an important role in many real-world applications, such as traffic forecasting, air pollutant forecasting, crowd-flow forecasting, and so on. State-of-the-art spatial-temporal forecasting models take…

机器学习 · 计算机科学 2024-01-22 Xinyu Su , Jianzhong Qi , Egemen Tanin , Yanchuan Chang , Majid Sarvi

Mapping a shape to some parametric domain is a fundamental tool in graphics and scientific computing. In practice, a map between two shapes is commonly represented by two meshes with same connectivity and different embedding. The standard…

计算几何 · 计算机科学 2020-12-16 Marco Livesu

We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward…

机器人学 · 计算机科学 2024-07-17 Jie Yin , Andrew Luo , Yilun Du , Anoop Cherian , Tim K. Marks , Jonathan Le Roux , Chuang Gan

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy…

图形学 · 计算机科学 2021-12-02 Seung Hyun Lee , Wonseok Roh , Wonmin Byeon , Sang Ho Yoon , Chan Young Kim , Jinkyu Kim , Sangpil Kim

Representation learning is a fundamental task in machine learning, aiming at uncovering structures from data to facilitate subsequent tasks. However, what is a good representation for planning and reasoning in a stochastic world remains an…

机器学习 · 计算机科学 2024-03-19 Meng Song

We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, and to associate them with a visual signal. Our method jointly…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Xixi Hu , Ziyang Chen , Andrew Owens

Urban sound has a huge influence over how we perceive places. Yet, city planning is concerned mainly with noise, simply because annoying sounds come to the attention of city officials in the form of complaints, while general urban sounds do…

社会与信息网络 · 计算机科学 2016-03-28 Luca Maria Aiello , Rossano Schifanella , Daniele Quercia , Francesco Aletta

Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object localization in videos.…

计算机视觉与模式识别 · 计算机科学 2021-11-11 Sizhe Li , Yapeng Tian , Chenliang Xu

This paper provides a formal and practical framework for sound abstraction of probabilistic actions. We start by precisely defining the concept of sound abstraction within the context of finite-horizon planning (where each plan is a finite…

人工智能 · 计算机科学 2013-02-18 AnHai Doan , Peter Haddawy

Non-line-of-sight localization in signal-deprived environments is a challenging yet pertinent problem. Acoustic methods in such predominantly indoor scenarios encounter difficulty due to the reverberant nature. In this study, we aim to…

机器学习 · 计算机科学 2024-04-03 Yi Di Yuan , Swee Liang Wong , Jonathan Pan

3D situational awareness is critical for any autonomous system. However, when operating underwater, environmental conditions often dictate the use of acoustic sensors. These acoustic sensors are plagued by high noise and a lack of 3D…

机器人学 · 计算机科学 2024-12-06 John McConnell , Ivana Collado-Gonzalez , Paul Szenher , Brendan Englot

Previous approaches in singer identification have used one of monophonic vocal tracks or mixed tracks containing multiple instruments, leaving a semantic gap between these two domains of audio. In this paper, we present a system to learn a…

声音 · 计算机科学 2019-06-27 Kyungyun Lee , Juhan Nam

In this paper, we describe a representation for spatial information, called the stochastic map, and associated procedures for building it, reading information from it, and revising it incrementally as new information is obtained. The map…

人工智能 · 计算机科学 2013-04-12 Randall Smith , Matthew Self , Peter Cheeseman

Simultaneous localization and mapping (SLAM) during communication is emerging. This technology promises to provide information on propagation environments and transceivers' location, thus creating several new services and applications for…

信息论 · 计算机科学 2021-08-10 Jie Yang , Chao-Kai Wen , Shi Jin , Xiao Li

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio…

声音 · 计算机科学 2019-05-15 Yu-Ding Lu , Hsin-Ying Lee , Hung-Yu Tseng , Ming-Hsuan Yang

We consider the sound ranging, or source localization, problem --- find the unknown source-point from known moments when the spherical wave of linearly, with time, increasing radius reaches known sensor-points --- in some non-proper metric…

泛函分析 · 数学 2019-11-01 Sergij V. Goncharov

The world of audio production and design has long been a difficult one to break into, requiring expertise and a working knowledge of the standard digital audio paradigms. This paper describes a novel interface that makes audio production…

人机交互 · 计算机科学 2020-10-01 Alexander Scarlatos