English
Related papers

Related papers: MV-UMI: A Scalable Multi-View Interface for Cross-…

200 papers

Reliable robotic grasping, especially with deformable objects such as fruits, remains a challenging task due to underactuated contact interactions with a gripper, unknown object dynamics and geometries. In this study, we propose a…

Robotics · Computer Science 2023-07-25 Yunhai Han , Kelin Yu , Rahul Batra , Nathan Boyd , Chaitanya Mehta , Tuo Zhao , Yu She , Seth Hutchinson , Ye Zhao

Imitation learning has emerged as a crucial ap proach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy generalization. However, existing methods often struggle to…

Robotics · Computer Science 2025-12-01 Yikai Tang , Haoran Geng , Sheng Zang , Pieter Abbeel , Jitendra Malik

Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting embodied robot interaction data remains prohibitively expensive. In contrast, human demonstration videos are far more…

Embodied learning for object-centric robotic manipulation is a rapidly developing and challenging area in embodied AI. It is crucial for advancing next-generation intelligent robots and has garnered significant interest recently. Unlike…

Robotics · Computer Science 2025-01-15 Ying Zheng , Lei Yao , Yuejiao Su , Yi Zhang , Yi Wang , Sicheng Zhao , Yiyi Zhang , Lap-Pui Chau

Grasping in cluttered scenes is challenging for robot vision systems, as detection accuracy can be hindered by partial occlusion of objects. We adopt a reinforcement learning (RL) framework and 3D vision architectures to search for feasible…

Robotics · Computer Science 2020-04-29 Xiangyu Chen , Zelin Ye , Jiankai Sun , Yuda Fan , Fang Hu , Chenxi Wang , Cewu Lu

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained…

Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view approaches often ignore cross-view relationships, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Xiangzuo Wu , Chengwei Ren , Jun Zhou , Xiu Li , Yuan Liu

The performance of image-based Reinforcement Learning (RL) agents can vary depending on the position of the camera used to capture the images. Training on multiple cameras simultaneously, including a first-person egocentric camera, can…

Machine Learning · Computer Science 2024-06-24 Mhairi Dunion , Stefano V. Albrecht

Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Han Wang , Jiayuan Zhang , Lipeng Wan , Xingyu Chen , Xuguang Lan , Nanning Zheng

Vision-based models for robotic grasping automate critical, repetitive, and draining industrial tasks. Existing approaches are typically limited in two ways: they either target a single gripper and are potentially applied on costly dual-arm…

Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Yihao Zhi , Chenghong Li , Hongjie Liao , Xihe Yang , Zhengwentai Sun , Jiahao Chang , Xiaodong Cun , Wensen Feng , Xiaoguang Han

Object grasping is an important ability required for various robot tasks. In particular, tasks that require precise force adjustments during operation, such as grasping an unknown object or using a grasped tool, are difficult for humans to…

Robotics · Computer Science 2024-01-22 Koki Yamane , Sho Sakaino , Toshiaki Tsuji

Collecting manipulation demonstrations with robotic hardware is tedious - and thus difficult to scale. Recording data on robot hardware ensures that it is in the appropriate format for Learning from Demonstrations (LfD) methods. By…

Robotics · Computer Science 2023-11-06 Kiran Doshi , Yijiang Huang , Stelian Coros

Masked Image Modeling (MIM) has garnered significant attention in self-supervised learning, thanks to its impressive capacity to learn scalable visual representations tailored for downstream tasks. However, images inherently contain…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Wenzhao Xiang , Chang Liu , Hongyang Yu , Xilin Chen

Within the imitation learning paradigm, training generalist robots requires large-scale datasets obtainable only through diverse curation. Due to the relative ease to collect, human demonstrations constitute a valuable addition when…

Robotics · Computer Science 2025-04-21 Yilong Song

Multi-perspective cameras are quickly gaining importance in many applications such as smart vehicles and virtual or augmented reality. However, a large system size or absence of overlap in neighbouring fields-of-view often complicate their…

Robotics · Computer Science 2022-04-14 Yifu Wang , Wenqing Jiang , Kun Huang , Sören Schwertfeger , Laurent Kneip

Dexterous hands exhibit significant potential for complex real-world grasping tasks. While recent studies have primarily focused on learning policies for specific robotic hands, the development of a universal policy that controls diverse…

Robotics · Computer Science 2024-10-04 Haoqi Yuan , Bohan Zhou , Yuhui Fu , Zongqing Lu

Embodied AI models often employ off the shelf vision backbones like CLIP to encode their visual observations. Although such general purpose representations encode rich syntactic and semantic information about the scene, much of this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Ainaz Eftekhar , Kuo-Hao Zeng , Jiafei Duan , Ali Farhadi , Ani Kembhavi , Ranjay Krishna

Embodied interaction has been introduced to human-robot interaction (HRI) as a type of teleoperation, in which users control robot arms with bodily action via handheld controllers or haptic gloves. Embodied teleoperation has made robot…

Robotics · Computer Science 2024-11-22 Siyou Pei , Alexander Chen , Ronak Kaoshik , Ruofei Du , Yang Zhang

Reconstructing real-world objects and estimating their movable joint structures are pivotal technologies within the field of robotics. Previous research has predominantly focused on supervised approaches, relying on extensively annotated…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Haowen Wang , Zhen Zhao , Zhao Jin , Zhengping Che , Liang Qiao , Yakun Huang , Zhipeng Fan , Xiuquan Qiao , Jian Tang