English
Related papers

Related papers: ROPA: Synthetic Robot Pose Generation for RGB-D Bi…

200 papers

We propose Human Pose Models that represent RGB and depth images of human poses independent of clothing textures, backgrounds, lighting conditions, body shapes and camera viewpoints. Learning such universal models requires training images…

Computer Vision and Pattern Recognition · Computer Science 2018-05-02 Jian Liu , Naveed Akhtar , Ajmal Mian

Despite the substantial progress in deep learning, its adoption in industrial robotics projects remains limited, primarily due to challenges in data acquisition and labeling. Previous sim2real approaches using domain randomization require…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Kaixin Bai , Lei Zhang , Zhaopeng Chen , Fang Wan , Jianwei Zhang

In recent times, object detection and pose estimation have gained significant attention in the context of robotic vision applications. Both the identification of objects of interest as well as the estimation of their pose remain important…

Robotics · Computer Science 2021-01-20 S. K. Paul , M. T. Chowdhury , M. Nicolescu , M. Nicolescu

Building generic robotic manipulation systems often requires large amounts of real-world data, which can be dificult to collect. Synthetic data generation offers a promising alternative, but limiting the sim-to-real gap requires significant…

Robotics · Computer Science 2024-11-18 Thomas Lips , Francis wyffels

Manipulation of deformable objects, such as ropes and cloth, is an important but challenging problem in robotics. We present a learning-based system where a robot takes as input a sequence of images of a human manipulating a rope from an…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Ashvin Nair , Dian Chen , Pulkit Agrawal , Phillip Isola , Pieter Abbeel , Jitendra Malik , Sergey Levine

Collecting manipulation demonstrations with robotic hardware is tedious - and thus difficult to scale. Recording data on robot hardware ensures that it is in the appropriate format for Learning from Demonstrations (LfD) methods. By…

Robotics · Computer Science 2023-11-06 Kiran Doshi , Yijiang Huang , Stelian Coros

Data labeling is a time intensive process. As such, many data scientists use various tools to aid in the data generation and labeling process. While these tools help automate labeling, many still require user interaction throughout the…

Machine Learning · Computer Science 2021-06-09 Kyle M. Hart , Ari B. Goodman , Ryan P. O'Shea

Accurate 6D object pose estimation is fundamental to robotic manipulation and grasping. Previous methods follow a local optimization approach which minimizes the distance between closest point pairs to handle the rotation ambiguity of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Meng Tian , Liang Pan , Marcelo H Ang , Gim Hee Lee

3D grasp synthesis generates grasping poses given an input object. Existing works tackle the problem by learning a direct mapping from objects to the distributions of grasping poses. However, because the physical contact is sensitive to…

Robotics · Computer Science 2023-05-09 Haoming Li , Xinzhuo Lin , Yang Zhou , Xiang Li , Yuchi Huo , Jiming Chen , Qi Ye

In this paper, we propose an iterative self-training framework for sim-to-real 6D object pose estimation to facilitate cost-effective robotic grasping. Given a bin-picking scenario, we establish a photo-realistic simulator to synthesize…

Robotics · Computer Science 2022-07-22 Kai Chen , Rui Cao , Stephen James , Yichuan Li , Yun-Hui Liu , Pieter Abbeel , Qi Dou

Robust object pose estimation is essential for manipulation and interaction tasks in robotics, particularly in scenarios where visual data is limited or sensitive to lighting, occlusions, and appearances. Tactile sensors often offer limited…

Video behavior recognition demands stable and discriminative representations under complex spatiotemporal variations. However, prevailing data augmentation strategies for videos remain largely perturbation-driven, often introducing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Feng-Qi Cui , Jinyang Huang , Sirui Zhao , Jinglong Guo , Qifan Cai , Xin Yan , Zhi Liu

Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise in this task,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Bao Li , Xiaomei Zhang , Miao Xu , Zhaoxin Fan , Xiangyu Zhu , Zhen Lei

Collaborative autonomous driving with multiple vehicles usually requires the data fusion from multiple modalities. To ensure effective fusion, the data from each individual modality shall maintain a reasonably high quality. However, in…

Artificial Intelligence · Computer Science 2024-08-02 Zhe Huang , Shuo Wang , Yongcai Wang , Wanting Li , Deying Li , Lei Wang

This paper proposes a method for hand pose estimation from RGB images that uses both external large-scale depth image datasets and paired depth and RGB images as privileged information at training time. We show that providing depth…

Computer Vision and Pattern Recognition · Computer Science 2018-11-20 Shanxin Yuan , Bjorn Stenger , Tae-Kyun Kim

6D pose estimation of textureless objects is a valuable but challenging task for many robotic applications. In this work, we propose a framework to address this challenge using only RGB images acquired from multiple viewpoints. The core…

Robotics · Computer Science 2023-02-23 Jun Yang , Wenjie Xue , Sahar Ghavidel , Steven L. Waslander

Concept personalization methods enable large text-to-image models to learn specific subjects (e.g., objects/poses/3D models) and synthesize renditions in new contexts. Given that the image references are highly biased towards visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 You Wu , Kean Liu , Xiaoyue Mi , Fan Tang , Juan Cao , Jintao Li

Estimating the 3D hand pose from a monocular RGB image is important but challenging. A solution is training on large-scale RGB hand images with accurate 3D hand keypoint annotations. However, it is too expensive in practice. Instead, we…

Computer Vision and Pattern Recognition · Computer Science 2020-10-06 Zhenyu Wu , Duc Hoang , Shih-Yao Lin , Yusheng Xie , Liangjian Chen , Yen-Yu Lin , Zhangyang Wang , Wei Fan

In this paper, we study the problem of adapting manipulation trajectories involving grasped objects (e.g. tools) defined for a single grasp pose to novel grasp poses. A common approach to address this is to define a new trajectory for each…

Robotics · Computer Science 2024-08-02 Georgios Papagiannis , Kamil Dreczkowski , Vitalis Vosylius , Edward Johns

Large-scale robot datasets have facilitated the learning of a wide range of robot manipulation skills, but these datasets remain difficult to collect and scale further, owing to the intractable amount of human time, effort, and cost…

Robotics · Computer Science 2026-03-27 Masoud Moghani , Mahdi Azizian , Animesh Garg , Yuke Zhu , Sean Huver , Ajay Mandlekar