English
Related papers

Related papers: InterACT: Inter-dependency Aware Action Chunking w…

200 papers

Cross-modal transfer learning is used to improve multi-modal classification models (e.g., for human activity recognition in human-robot collaboration). However, existing methods require paired sensor data at both training and inference,…

Machine Learning · Computer Science 2025-09-15 Leen Daher , Zhaobo Wang , Malcolm Mielle

Human activity recognition in videos has been widely studied and has recently gained significant advances with deep learning approaches; however, it remains a challenging task. In this paper, we propose a novel framework that simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2021-01-25 Dong-Gyu Lee , Seong-Whan Lee

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

Bimanual object manipulation involves multiple visuo-haptic sensory feedbacks arising from the interaction with the environment that are managed from the central nervous system and consequently translated in motor commands. Kinematic…

Robotics · Computer Science 2022-11-23 Elisa Galofaro , Erika D'Antonio , Nicola Lotti , Fabrizio Patane' , Maura Casadio , Lorenzo Masia

With the goal of increasing the speed and efficiency in robotic dual arm manipulation, a novel control approach is presented that utilizes intentional simultaneous impacts to rapidly grasp objects. This approach uses the time-invariant…

Robotics · Computer Science 2023-04-25 Jari J. van Steen , Abdullah Coşgun , Nathan van de Wouw , Alessandro Saccon

Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these…

Robotics · Computer Science 2023-04-27 Tony Z. Zhao , Vikash Kumar , Sergey Levine , Chelsea Finn

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements.…

Imitation Learning (IL) is a powerful paradigm to teach robots to perform manipulation tasks by allowing them to learn from human demonstrations collected via teleoperation, but has mostly been limited to single-arm manipulation. However,…

Learning bimanual manipulation is challenging due to its high dimensionality and tight coordination required between two arms. Eye-in-hand imitation learning, which uses wrist-mounted cameras, simplifies perception by focusing on…

Robotics · Computer Science 2025-08-19 I-Chun Arthur Liu , Jason Chen , Gaurav Sukhatme , Daniel Seita

Robotic manipulation is essential for the widespread adoption of robots in industrial and home settings and has long been a focus within the robotics community. Advances in artificial intelligence have introduced promising learning-based…

Robotics · Computer Science 2025-03-04 Kelin Li , Shubham M Wagh , Nitish Sharma , Saksham Bhadani , Wei Chen , Chang Liu , Petar Kormushev

We address the problem of safely solving complex bimanual robot manipulation tasks with sparse rewards. Such challenging tasks can be decomposed into sub-tasks that are accomplishable by different robots concurrently or sequentially for…

Machine Learning · Computer Science 2021-10-07 Minghao Zhang , Pingcheng Jian , Yi Wu , Huazhe Xu , Xiaolong Wang

Bimanual manipulation, fundamental to human daily activities, remains a challenging task due to its inherent complexity of coordinated control. Recent advances have enabled zero-shot learning of single-arm manipulation skills through…

Robotics · Computer Science 2025-07-29 Ziyin Xiong , Yinghan Chen , Puhao Li , Yixin Zhu , Tengyu Liu , Siyuan Huang

To fluently collaborate with people, robots need the ability to recognize human activities accurately. Although modern robots are equipped with various sensors, robust human activity recognition (HAR) still remains a challenging task for…

Robotics · Computer Science 2020-08-17 Md Mofijul Islam , Tariq Iqbal

We present a novel hierarchical spatiotemporal action tokenizer for in-context imitation learning. We first propose a hierarchical approach, which consists of two successive levels of vector quantization. In particular, the lower level…

The true promise of humanoid robotics lies beyond single-agent autonomy: two or more humanoids must engage in physically grounded, socially meaningful whole-body interactions that echo the richness of human social interaction. However,…

Robotics · Computer Science 2025-10-14 Zuhong Liu , Junhao Ge , Minhao Xiong , Jiahao Gu , Bowei Tang , Wei Jing , Siheng Chen

Bimanual manipulation tasks typically involve multiple stages which require efficient interactions between two arms, posing step-wise and stage-wise challenges for imitation learning systems. Specifically, failure and delay of one step will…

Robotics · Computer Science 2024-09-05 Dongjie Yu , Hang Xu , Yizhou Chen , Yi Ren , Jia Pan

Humans throw and catch objects all the time. However, such a seemingly common skill introduces a lot of challenges for robots to achieve: The robots need to operate such dynamic actions at high-speed, collaborate precisely, and interact…

Robotics · Computer Science 2023-09-12 Binghao Huang , Yuanpei Chen , Tianyu Wang , Yuzhe Qin , Yaodong Yang , Nikolay Atanasov , Xiaolong Wang

Underwater robotic manipulation remains challenging because lighting variation, color attenuation, scattering, and reduced visibility can severely degrade visuomotor policies. We present Bi-AQUA, the first underwater bilateral control-based…

Robotics · Computer Science 2026-03-09 Takeru Tsunoori , Masato Kobayashi , Yuki Uranishi

We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Jihyun Lee , Minhyuk Sung , Honggyu Choi , Tae-Kyun Kim

Recent progress has been made in using attention based encoder-decoder framework for image and video captioning. Most existing decoders apply the attention mechanism to every generated word including both visual words (e.g., "gun" and…

Computer Vision and Pattern Recognition · Computer Science 2018-12-31 Jingkuan Song , Xiangpeng Li , Lianli Gao , Heng Tao Shen