English
Related papers

Related papers: Kaiwu: A Multimodal Manipulation Dataset and Frame…

200 papers

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Imitation learning from human demonstrations is an effective means to teach robots manipulation skills. But data acquisition is a major bottleneck in applying this paradigm more broadly, due to the amount of cost and human effort involved.…

Robotics · Computer Science 2025-03-07 Zhenyu Jiang , Yuqi Xie , Kevin Lin , Zhenjia Xu , Weikang Wan , Ajay Mandlekar , Linxi Fan , Yuke Zhu

A major challenge for the realization of intelligent robots is to supply them with cognitive abilities in order to allow ordinary users to program them easily and intuitively. One way of such programming is teaching work tasks by…

Human-Computer Interaction · Computer Science 2016-11-17 P. C. McGuire , J. Fritsch , J. J. Steil , F. Roethling , G. A. Fink , S. Wachsmuth , G. Sagerer , H. Ritter

Recently, instruction-following audio-language models have received broad attention for audio interaction with humans. However, the absence of pre-trained audio models capable of handling diverse audio types and tasks has hindered progress…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Yunfei Chu , Jin Xu , Xiaohuan Zhou , Qian Yang , Shiliang Zhang , Zhijie Yan , Chang Zhou , Jingren Zhou

We introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety of biology experiments, and includes triplets of interacting…

Imitating human demonstrations is a promising approach to endow robots with various manipulation capabilities. While recent advances have been made in imitation learning and batch (offline) reinforcement learning, a lack of open-source…

Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual contents. Existing state-of-the-art approaches use…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Zhou Yu , Yuhao Cui , Jun Yu , Dacheng Tao , Qi Tian

For safe and effective operation of humanoid robots in human-populated environments, the problem of commanding a large number of Degrees of Freedom (DoF) while simultaneously considering dynamic obstacles and human proximity has still not…

Robotics · Computer Science 2024-12-12 Jakub Rozlivek , Alessandro Roncone , Ugo Pattacini , Matej Hoffmann

Information processing tasks involve complex cognitive mechanisms that are shaped by various factors, including individual goals, prior experience, and system environments. Understanding such behaviors requires a sophisticated and…

Human-Computer Interaction · Computer Science 2025-07-24 Kaixin Ji , Danula Hettiachchi , Falk Scholer , Flora D. Salim , Damiano Spina

The predictive functions that permit humans to infer their body state by sensorimotor integration are critical to perform safe interaction in complex environments. These functions are adaptive and robust to non-linear actuators and noisy…

Robotics · Computer Science 2019-10-24 Pablo Lanillos , Gordon Cheng

To enhance human-robot social interaction, it is essential for robots to process multiple social cues in a complex real-world environment. However, incongruency of input information across modalities is inevitable and could be challenging…

Robotics · Computer Science 2023-03-14 Di Fu , Fares Abawi , Hugo Carneiro , Matthias Kerzel , Ziwei Chen , Erik Strahl , Xun Liu , Stefan Wermter

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Intelligent instruction-following robots capable of improving from autonomously collected experience have the potential to transform robot learning: instead of collecting costly teleoperated demonstration data, large-scale deployment of…

Robotics · Computer Science 2025-02-26 Zhiyuan Zhou , Pranav Atreya , Abraham Lee , Homer Walke , Oier Mees , Sergey Levine

The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Haoxin Lv , Ijazul Haq , Jin Du , Jiaxin Ma , Binnian Zhu , Xiaobing Dang , Chaoan Liang , Ruxu Du , Yingjie Zhang , Muhammad Saqib

We report on an extensive study of the benefits and limitations of current deep learning approaches to object recognition in robot vision scenarios, introducing a novel dataset used for our investigation. To avoid the biases in currently…

Robotics · Computer Science 2021-08-25 Giulia Pasquale , Carlo Ciliberto , Francesca Odone , Lorenzo Rosasco , Lorenzo Natale

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Bizhu Wu , Jinheng Xie , Meidan Ding , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Large behavior models have shown strong dexterous manipulation capabilities by extending imitation learning to large-scale training on multi-task robot data, yet their generalization remains limited by the insufficient robot data coverage.…

Physical human-robot interaction can improve human ergonomics, task efficiency, and the flexibility of automation, but often requires application-specific methods to detect human state and determine robot response. At the same time, many…

Robotics · Computer Science 2026-02-17 Kevin Haninger , Christian Hegeler , Luka Peternel

The integration of dialogue interfaces in mobile devices has become ubiquitous, providing a wide array of services. As technology progresses, humanoid robots designed with human-like features to interact effectively with people are gaining…

Human-Computer Interaction · Computer Science 2025-01-14 Mohamed Ala Yahyaoui , Mouaad Oujabour , Leila Ben Letaifa , Amine Bohi

Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often learn the kinematic structure of the robot implicitly and face challenges due to the…

Robotics · Computer Science 2026-04-17 Roman Freiberg , Alexander Qualmann , Ngo Anh Vien , Gerhard Neumann