English
Related papers

Related papers: MOSAIC: Learning Unified Multi-Sensory Object Prop…

200 papers

Inferring the unseen attribute-object composition is critical to make machines learn to decompose and compose complex concepts like people. Most existing methods are limited to the composition recognition of single-attribute-object, and can…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Hui Chen , Jingjing Jiang , Nanning Zheng

Nowadays, robots become a companion in everyday life. To be well-accepted by humans, robots should efficiently understand meanings of their partners' motions and body language, and respond accordingly. Learning concepts by imitation brings…

Artificial Intelligence · Computer Science 2017-07-25 Mina Alibeigi , Majid Nili Ahmadabadi , Babak Nadjar Araabi

Humans have impressive generalization capabilities when it comes to manipulating objects and tools in completely novel environments. These capabilities are, at least partially, a result of humans having internal models of their bodies and…

Robotics · Computer Science 2021-06-28 Sarah Bechtle , Neha Das , Franziska Meier

Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to…

Robotics · Computer Science 2024-10-14 Irving Fang , Kairui Shi , Xujin He , Siqi Tan , Yifan Wang , Hanwen Zhao , Hung-Jui Huang , Wenzhen Yuan , Chen Feng , Jing Zhang

Tactile sensing is critical to fine-grained, contact-rich manipulation tasks, such as insertion and assembly. Prior research has shown the possibility of learning tactile-guided policy from teleoperated demonstration data. However, to…

Robotics · Computer Science 2025-02-07 Kelin Yu , Yunhai Han , Qixian Wang , Vaibhav Saxena , Danfei Xu , Ye Zhao

Humans interact in rich and diverse ways with the environment. However, the representation of such behavior by artificial agents is often limited. In this work we present \textit{motion concepts}, a novel multimodal representation of human…

Computer Vision and Pattern Recognition · Computer Science 2019-03-07 Miguel Vasco , Francisco S. Melo , David Martins de Matos , Ana Paiva , Tetsunari Inamura

Deep-learning and large scale language-image training have produced image object detectors that generalise well to diverse environments and semantic classes. However, single-image object detectors trained on internet data are not optimally…

Robotics · Computer Science 2024-02-07 Nicolas Harvey Chapman , Feras Dayoub , Will Browne , Chris Lehnert

This paper proposes a novel framework for utilizing skin sensors as a new operation interface of complex robots. The skin sensors employed in this study possess the capability to quantify multimodal tactile information at multiple contact…

The field of robotic manipulation has advanced significantly in recent years. At the sensing level, several novel tactile sensors have been developed, capable of providing accurate contact information. On a methodological level, learning…

Robotics · Computer Science 2026-04-21 Niklas Funk , Changqi Chen , Tim Schneider , Georgia Chalvatzaki , Roberto Calandra , Jan Peters

Conventional object detection models require large amounts of training data. In comparison, humans can recognize previously unseen objects by merely knowing their semantic description. To mimic similar behaviour, zero-shot object detection…

Computer Vision and Pattern Recognition · Computer Science 2020-04-03 Shafin Rahman , Salman Khan , Nick Barnes

Human-robot collaboration requires the contactless estimation of the physical properties of containers manipulated by a person, for example while pouring content in a cup or moving a food box. Acoustic and visual signals can be used to…

Multimedia · Computer Science 2022-03-07 A. Xompero , Y. L. Pang , T. Patten , A. Prabhakar , B. Calli , A. Cavallaro

The Platonic Representation Hypothesis claims that recent foundation models are converging to a shared representation space as a function of their downstream task performance, irrespective of the objectives and data modalities used to train…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Laure Ciernik , Lorenz Linhardt , Marco Morik , Jonas Dippel , Simon Kornblith , Lukas Muttenthaler

Tactile and visual perception are both crucial for humans to perform fine-grained interactions with their environment. Developing similar multi-modal sensing capabilities for robots can significantly enhance and expand their manipulation…

Robotics · Computer Science 2025-01-08 Binghao Huang , Yixuan Wang , Xinyi Yang , Yiyue Luo , Yunzhu Li

Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensionality and…

To be useful in everyday environments, robots must be able to observe and learn about objects. Recent datasets enable progress for classifying data into known object categories; however, it is unclear how to collect reliable object data…

Robotics · Computer Science 2019-01-18 Abhishek Venkataraman , Brent Griffin , Jason J. Corso

Current motion-based multiple object tracking (MOT) approaches rely heavily on Intersection-over-Union (IoU) for object association. Without using 3D features, they are ineffective in scenarios with occlusions or visually similar objects.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Milad Khanchi , Maria Amer , Charalambos Poullis

The integration of visual-tactile stimulus is common while humans performing daily tasks. In contrast, using unimodal visual or tactile perception limits the perceivable dimensionality of a subject. However, it remains a challenge to…

Robotics · Computer Science 2019-02-19 Jet-Tsyn Lee , Danushka Bollegala , Shan Luo

This paper proposes a novel contrastive learning framework, called FOCAL, for extracting comprehensive features from multimodal time-series sensing signals through self-supervised training. Existing multimodal contrastive frameworks mostly…

Artificial Intelligence · Computer Science 2023-11-01 Shengzhong Liu , Tomoyoshi Kimura , Dongxin Liu , Ruijie Wang , Jinyang Li , Suhas Diggavi , Mani Srivastava , Tarek Abdelzaher

Combined visual and force feedback play an essential role in contact-rich robotic manipulation tasks. Current methods focus on developing the feedback control around a single modality while underrating the synergy of the sensors. Fusing…

Robotics · Computer Science 2022-02-18 Piaopiao Jin , Yinjie Lin , Yanchao Tan , Tiefeng Li , Wei Yang

The goal of multi-object tracking (MOT) is to detect and track all objects in a scene across frames, while maintaining a unique identity for each object. Most existing methods rely on the spatial-temporal motion features and appearance…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Yanzhao Fang