中文
相关论文

相关论文: Swoosh! Rattle! Thump! -- Actions that Sound

200 篇论文

Interacting with bins and containers is a fundamental task in robotics, making state estimation of the objects inside the bin critical. While robots often use cameras for state estimation, the visual modality is not always ideal due to…

计算机视觉与模式识别 · 计算机科学 2021-10-26 Boyuan Chen , Mia Chiquier , Hod Lipson , Carl Vondrick

Humans and other intelligent animals evolved highly sophisticated perception systems that combine multiple sensory modalities. On the other hand, state-of-the-art artificial agents rely mostly on visual inputs or structured low-dimensional…

机器学习 · 计算机科学 2021-07-07 Shashank Hegde , Anssi Kanervisto , Aleksei Petrenko

Although pre-training on a large amount of data is beneficial for robot learning, current paradigms only perform large-scale pretraining for visual representations, whereas representations for other modalities are trained from scratch. In…

机器人学 · 计算机科学 2024-05-15 Jared Mejia , Victoria Dean , Tess Hellebrekers , Abhinav Gupta

Humans use all of their senses to accomplish different tasks in everyday activities. In contrast, existing work on robotic manipulation mostly relies on one, or occasionally two modalities, such as vision and touch. In this work, we…

机器人学 · 计算机科学 2022-12-09 Hao Li , Yizhi Zhang , Junzhe Zhu , Shaoxiong Wang , Michelle A Lee , Huazhe Xu , Edward Adelson , Li Fei-Fei , Ruohan Gao , Jiajun Wu

A key challenge in robotic food manipulation is modeling the material properties of diverse and deformable food items. We propose using a multimodal sensory approach to interact and play with food that facilitates the ability to distinguish…

机器人学 · 计算机科学 2021-01-08 Amrita Sawhney , Steven Lee , Kevin Zhang , Manuela Veloso , Oliver Kroemer

Robotic failure is all too common in unstructured robot tasks. Despite well designed controllers, robots often fail due to unexpected events. How do robots measure unexpected events? Many do not. Most robots are driven by the senseplan- act…

机器人学 · 计算机科学 2016-09-19 Juan Rojas , Zhengjie Huang , Shuangqi Luo , Yunlong Du Wenwei Kuang , Dingqiao Zhu , Kensuke Harada

Audio perception is a key to solving a variety of problems ranging from acoustic scene analysis, music meta-data extraction, recommendation, synthesis and analysis. It can potentially also augment computers in doing tasks that humans do…

声音 · 计算机科学 2020-02-12 Prateek Verma , Kenneth Salisbury

Perception in robot manipulation has been actively explored with the goal of advancing and integrating vision and touch for global and local feature extraction. However, it is difficult to perceive certain object internal states, and the…

机器人学 · 计算机科学 2023-08-04 Shihan Lu , Heather Culbertson

We present a novel approach for interactive auditory object analysis with a humanoid robot. The robot elicits sensory information by physically shaking visually indistinguishable plastic capsules. It gathers the resulting audio signals from…

机器人学 · 计算机科学 2018-07-11 Manfred Eppe , Matthias Kerzel , Erik Strahl , Stefan Wermter

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-verbal gestures, facial expressions, and body language. In this…

机器学习 · 计算机科学 2026-05-13 Alice Zhang , Callihan Bertley , Dawei Liang , Edison Thomaz

Humans learn about objects via interaction and using multiple perceptions, such as vision, sound, and touch. While vision can provide information about an object's appearance, non-visual sensors, such as audio and haptics, can provide…

机器人学 · 计算机科学 2023-09-18 Gyan Tatiya , Jonathan Francis , Jivko Sinapov

We aim for zero-shot localization and classification of human actions in video. Where traditional approaches rely on global attribute or object classification scores for their zero-shot knowledge transfer, our main contribution is a…

计算机视觉与模式识别 · 计算机科学 2017-12-14 Pascal Mettes , Cees G. M. Snoek

As human-robot collaboration is becoming more widespread, there is a need for a more natural way of communicating with the robot. This includes combining data from several modalities together with the context of the situation and background…

人机交互 · 计算机科学 2024-04-03 Petr Vanc , Radoslav Skoviera , Karla Stepanova

Sounds originate from object motions and vibrations of surrounding air. Inspired by the fact that humans is capable of interpreting sound sources from how objects move visually, we propose a novel system that explicitly captures such motion…

计算机视觉与模式识别 · 计算机科学 2019-04-16 Hang Zhao , Chuang Gan , Wei-Chiu Ma , Antonio Torralba

We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Yiming Dou , Wonseok Oh , Yuqing Luo , Antonio Loquercio , Andrew Owens

Human action is naturally compositional: humans can easily recognize and perform actions with objects that are different from those used in training demonstrations. In this paper, we study the compositionality of action by looking into the…

计算机视觉与模式识别 · 计算机科学 2020-09-15 Joanna Materzynska , Tete Xiao , Roei Herzig , Huijuan Xu , Xiaolong Wang , Trevor Darrell

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label…

声音 · 计算机科学 2025-07-17 Jaesung Huh , Jacob Chalk , Evangelos Kazakos , Dima Damen , Andrew Zisserman

Recent Large Audio-Language Models (LALMs) have shown strong performance on various audio understanding tasks such as speech translation and Audio Q\&A. However, they exhibit significant limitations on challenging audio reasoning tasks in…

计算与语言 · 计算机科学 2025-09-29 Zhen Xiong , Yujun Cai , Zhecheng Li , Junsong Yuan , Yiwei Wang

We propose Audio Noise Awareness using Visuals of Indoors for NAVIgation for quieter robot path planning. While humans are naturally aware of the noise they make and its impact on those around them, robots currently lack this awareness. A…

机器人学 · 计算机科学 2024-10-25 Vidhi Jain , Rishi Veerapaneni , Yonatan Bisk

Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "You Only Speak Once to…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Lei Li