English
Related papers

Related papers: DSPv2: Improved Dense Policy for Effective and Gen…

200 papers

Acquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Junjie Zhang , Chenjia Bai , Haoran He , Wenke Xia , Zhigang Wang , Bin Zhao , Xiu Li , Xuelong Li

Developing personal robots that can perform a diverse range of manipulation tasks in unstructured environments necessitates solving several challenges for robotic grasping systems. We take a step towards this broader goal by presenting the…

We present AnchorDP3, a diffusion policy framework for dual-arm robotic manipulation that achieves state-of-the-art performance in highly randomized environments. AnchorDP3 integrates three key innovations: (1) Simulator-Supervised Semantic…

Robotics · Computer Science 2025-06-26 Ziyan Zhao , Ke Fan , He-Yang Xu , Ning Qiao , Bo Peng , Wenlong Gao , Dongjiang Li , Hui Shen

Imitation learning has emerged as a crucial ap proach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy generalization. However, existing methods often struggle to…

Robotics · Computer Science 2025-12-01 Yikai Tang , Haoran Geng , Sheng Zang , Pieter Abbeel , Jitendra Malik

Policy search methods can allow robots to learn control policies for a wide range of tasks, but practical applications of policy search often require hand-engineered components for perception, state estimation, and low-level control. In…

Machine Learning · Computer Science 2016-04-20 Sergey Levine , Chelsea Finn , Trevor Darrell , Pieter Abbeel

Recently, 3D vision-based diffusion policies have shown strong capability in learning complex robotic manipulation skills. However, a common architectural mismatch exists in these models: a tiny yet efficient point-cloud encoder is often…

Robotics · Computer Science 2026-02-02 Jinhao Zhang , Zhexuan Zhou , Huizhe Li , Yichen Lai , Wenlong Xia , Haoming Song , Youmin Gong , Jie Mei

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches…

Robotics · Computer Science 2025-09-19 Anzhe Chen , Yifei Yang , Zhenjie Zhu , Kechun Xu , Zhongxiang Zhou , Rong Xiong , Yue Wang

We address dynamic manipulation of deformable linear objects by presenting SPiD, a physics-informed self-supervised learning framework that couples an accurate deformable object model with an augmented self-supervised training strategy. On…

Robotics · Computer Science 2026-02-04 Youyuan Long , Gokhan Solak , Sara Zeynalpour , Heng Zhang , Arash Ajoudani

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Yang Xing , Jiong Wu , Savas Ozdemir , Ying Zhang , Yang Yang , Wei Shao , Kuang Gong

Diffusion policies have recently emerged as a powerful class of visuomotor controllers for robot manipulation, offering stable training and expressive multi-modal action modeling. However, existing approaches typically treat action…

Robotics · Computer Science 2025-10-01 Zezeng Li , Rui Yang , Ruochen Chen , ZhongXuan Luo , Liming Chen

Humanoid robot manipulation is a crucial research area for executing diverse human-level tasks, involving high-level semantic reasoning and low-level action generation. However, precise scene understanding and sample-efficient learning from…

Robotics · Computer Science 2026-01-15 Xuetao Li , Wenke Huang , Mang Ye , Jifeng Xuan , Bo Du , Sheng Liu , Miao Li

Vision-based policies are widely applied in robotics for tasks such as manipulation and locomotion. On lightweight mobile robots, however, they face a trilemma of limited scene transferability, restricted onboard computation resources, and…

Robotics · Computer Science 2026-03-24 Kai Li , Shiyu Zhao

Is it possible to learn policies for robotic assembly that can generalize to new objects? We explore this idea in the context of the kit assembly task. Since classic methods rely heavily on object pose estimation, they often struggle to…

Robotics · Computer Science 2020-05-19 Kevin Zakka , Andy Zeng , Johnny Lee , Shuran Song

As robots are increasingly deployed in diverse application domains, enabling robust mobility across different embodiments has become a critical challenge. Classical mobility stacks, though effective on specific platforms, require extensive…

Robotics · Computer Science 2025-10-29 Wei Liu , Huihua Zhao , Chenran Li , Yuchen Deng , Joydeep Biswas , Soha Pouya , Yan Chang

What is the right object representation for manipulation? We would like robots to visually perceive scenes and learn an understanding of the objects in them that (i) is task-agnostic and can be used as a building block for a variety of…

Robotics · Computer Science 2018-09-10 Peter R. Florence , Lucas Manuelli , Russ Tedrake

Deep imitation learning is promising for robot manipulation because it only requires demonstration samples. In this study, deep imitation learning is applied to tasks that require force feedback. However, existing demonstration methods have…

Robotics · Computer Science 2024-02-27 Heecheol Kim , Yoshiyuki Ohmura , Akihiko Nagakubo , Yasuo Kuniyoshi

Classical Visual Servoing (VS) rely on handcrafted visual features, which limit their generalizability. Recently, a number of approaches, some based on Deep Neural Networks, have been proposed to overcome this limitation by comparing…

Robotics · Computer Science 2022-01-21 Nicholas Adrian , Van-Thach Do , Quang-Cuong Pham

Contact force in contact-rich environments is an essential modality for robots to perform general-purpose manipulation tasks, as it provides information to compensate for the deficiencies of visual and proprioceptive data in collision…

Robotics · Computer Science 2024-11-13 Bo Zhou , Ruixuan Jiao , Yi Li , Xiaogang Yuan , Fang Fang , Shihua Li

Agile control of mobile manipulator is challenging because of the high complexity coupled by the robotic system and the unstructured working environment. Tracking and grasping a dynamic object with a random trajectory is even harder. In…

Robotics · Computer Science 2020-06-09 Cong Wang , Qifeng Zhang , Qiyan Tian , Shuo Li , Xiaohui Wang , David Lane , Yvan Petillot , Ziyang Hong , Sen Wang

Manipulating clusters of deformable objects presents a substantial challenge with widespread applicability, but requires contact-rich whole-arm interactions. A potential solution must address the limited capacity for realistic model…