English
Related papers

Related papers: Manipulation as in Simulation: Enabling Accurate G…

200 papers

Human body motions can be captured as a high-dimensional continuous signal using motion sensor technologies. The resulting data can be surprisingly rich in information, even when captured from persons with limited mobility. In this work, we…

Keypoint detection is an essential building block for many robotic applications like motion capture and pose estimation. Historically, keypoints are detected using uniquely engineered markers such as checkerboards or fiducials. More…

Robotics · Computer Science 2023-02-28 Jingpei Lu , Florian Richter , Michael Yip

A vast literature shows that the learning-based visual perception model is sensitive to adversarial noises, but few works consider the robustness of robotic perception models under widely-existing camera motion perturbations. To this end,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Hanjiang Hu , Zuxin Liu , Linyi Li , Jiacheng Zhu , Ding Zhao

Deep learning has provided new ways of manipulating, processing and analyzing data. It sometimes may achieve results comparable to, or surpassing human expert performance, and has become a source of inspiration in the era of artificial…

Providing mobile robots with the ability to manipulate objects has, despite decades of research, remained a challenging problem. The problem is approachable in constrained environments where there is ample prior knowledge of the environment…

Robotics · Computer Science 2022-06-08 David Watkins

Accurate deformable object manipulation (DOM) is essential for achieving autonomy in robotic surgery, where soft tissues are being displaced, stretched, and dissected. Many DOM methods can be powered by simulation, which ensures realistic…

Robotics · Computer Science 2024-05-31 Xiao Liang , Fei Liu , Yutong Zhang , Yuelei Li , Shan Lin , Michael Yip

Latent diffusion models (LDMs) exhibit an impressive ability to produce realistic images, yet the inner workings of these models remain mysterious. Even when trained purely on images without explicit depth information, they typically output…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Yida Chen , Fernanda Viégas , Martin Wattenberg

Contemporary interventional imaging lacks the real-time 3D guidance needed for the precise localization of mobile thoracic targets. While Cone-Beam CT (CBCT) provides 3D data, it is often too slow for dynamic motion tracking. Deep learning…

Medical Physics · Physics 2025-11-19 Fawazilla Utomo , Tess Reynolds , Nicholas Hindley

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Yucheng Hu , Yanjiang Guo , Pengchao Wang , Xiaoyu Chen , Yen-Jen Wang , Jianke Zhang , Koushil Sreenath , Chaochao Lu , Jianyu Chen

This study seeks to automate camera movement control for filming existing subjects into attractive videos, contrasting with the creation of non-existent content by directly generating the pixels. We select drone videos as our test case due…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Yunzhong Hou , Liang Zheng , Philip Torr

The incorporation of world modeling into manipulation policy learning has pushed the boundary of manipulation performance. However, existing efforts simply model the 2D visual dynamics, which is insufficient for robust manipulation when…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yuxin He , Ruihao Zhang , Xianzu Wu , Zhiyuan Zhang , Cheng Ding , Qiang Nie

Robust imitation learning for robot manipulation requires comprehensive 3D perception, yet many existing methods struggle in cluttered environments. Fixed camera view approaches are vulnerable to perspective changes, and 3D point cloud…

Robotics · Computer Science 2025-07-08 Daqi Huang , Zhehao Cai , Yuzhi Hao , Zechen Li , Chee-Meng Chew

Navigating unknown environments with a single RGB camera is challenging, as the lack of depth information prevents reliable collision-checking. While some methods use estimated depth to build collision maps, we found that depth estimates…

Robotics · Computer Science 2025-11-27 Basant Sharma , Prajyot Jadhav , Pranjal Paul , K. Madhava Krishna , Arun Kumar Singh

Harnessing human movements to command an Unmanned Aerial Vehicle (UAV) holds the potential to revolutionize their deployment, rendering it more intuitive and user-centric. In this research, we introduce a novel methodology adept at…

Robotics · Computer Science 2024-08-20 Akash Chaudhary , Tiago Nascimento , Martin Saska

Recent progress in human shape learning, shows that neural implicit models are effective in generating 3D human surfaces from limited number of views, and even from a single RGB image. However, existing monocular approaches still struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Marco Pesavento , Yuanlu Xu , Nikolaos Sarafianos , Robert Maier , Ziyan Wang , Chun-Han Yao , Marco Volino , Edmond Boyer , Adrian Hilton , Tony Tung

We consider the problem of robotic grasping using depth + RGB information sampling from a real sensor. we design an encoder-decoder neural network to predict grasp policy in real time. This method can fuse the advantage of depth image and…

Robotics · Computer Science 2019-06-03 Song Yaoxian , Cheng Chun , Fei Yuejiao , Li Xiangqing , Yu Changbin

Depth cameras are frequently used in robotic manipulation, e.g. for visual servoing. The quality of small and compact depth cameras is though often not sufficient for depth reconstruction, which is required for precise tracking in and…

Machine Learning · Computer Science 2023-05-11 Claudius Kienle , David Petri

We consider the problem of segmenting dynamic regions in CrowdCam images, where a dynamic region is the projection of a moving 3D object on the image plane. Quite often, these regions are the most interesting parts of an image. CrowdCam…

Computer Vision and Pattern Recognition · Computer Science 2019-06-25 Nir Zarrabi , Shai Avidan , Yael Moses

Careful robot manipulation in every-day cluttered environments requires an accurate understanding of the 3D scene, in order to grasp and place objects stably and reliably and to avoid colliding with other objects. In general, we must…

Robotics · Computer Science 2025-11-11 Aditya Agarwal , Gaurav Singh , Bipasha Sen , Tomás Lozano-Pérez , Leslie Pack Kaelbling

Mobile robots require accurate and robust depth measurements to understand and interact with the environment. While existing sensing modalities address this problem to some extent, recent research on monocular depth estimation has leveraged…

Robotics · Computer Science 2024-10-02 Marco Job , Thomas Stastny , Tim Kazik , Roland Siegwart , Michael Pantic