English
Related papers

Related papers: Vid2Sid: Videos Can Help Close the Sim2Real Gap

200 papers

Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Muhua Zhu , Xinhao Jin , Yu Zhang , Yifei Xue , Tie Ji , Yizhen Lao

Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies…

Robotics · Computer Science 2025-06-23 Kaiyuan Chen , Shuangyu Xie , Zehan Ma , Pannag R Sanketi , Ken Goldberg

Constructing physically accurate simulation environments (Real2Sim) traditionally relies on manual system identification or rigid, exhaustive exploration routines. These task-agnostic pipelines often fail to leverage semantic scene context,…

Robotics · Computer Science 2026-05-19 Alessandro Adami , Sebastian Zudaire , Ruggero Carli , Pietro Falco

Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated data, the domain gap between synthetic and real-world driving videos significantly limits…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Haonan Zhao , Yiting Wang , Jingkun Chen , Valentina Donzella , Thomas Bashford-Rogers , Kurt Debattista

High-fidelity physics simulation is essential for scalable robotic learning, but the sim-to-real gap persists, especially for tasks involving complex, dynamic, and discontinuous interactions like physical contacts. Explicit system…

Robotics · Computer Science 2026-01-21 Changwei Jing , Jai Krishna Bandi , Jianglong Ye , Yan Duan , Pieter Abbeel , Xiaolong Wang , Sha Yi

Visual Odometry (VO) is essential to downstream mobile robotics and augmented/virtual reality tasks. Despite recent advances, existing VO methods still rely on heuristic design choices that require several weeks of hyperparameter tuning by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Nico Messikommer , Giovanni Cioffi , Mathias Gehrig , Davide Scaramuzza

Most current works in Sim2Real learning for robotic manipulation tasks leverage camera vision that may be significantly occluded by robot hands during the manipulation. Tactile sensing offers complementary information to vision and can…

Robotics · Computer Science 2021-01-19 Daniel Fernandes Gomes , Paolo Paoletti , Shan Luo

Robotic assistance allows surgeries to be reliably and accurately executed while still under direct supervision of the surgeon, combining the strengths of robotic technology with the surgeon's expertise. This paper describes a robotic…

Systems and Control · Electrical Eng. & Systems 2025-12-09 Daniel Larby , Joshua Kershaw , Matthew Allen , Fulvio Forni

Laparoscopic surgery constrains surgeons spatial awareness because procedures are performed through a monocular, two-dimensional (2D) endoscopic view. Conventional training methods using dry-lab models or recorded videos provide limited…

Human-Computer Interaction · Computer Science 2025-11-05 Songyang Liu , Yunpeng Tan , Shuai Li

Accurately modeling soft robots in simulation is computationally expensive and commonly falls short of representing the real world. This well-known discrepancy, known as the sim-to-real gap, can have several causes, such as coarsely…

Robotics · Computer Science 2024-09-10 Junpeng Gao , Mike Yan Michelis , Andrew Spielberg , Robert K. Katzschmann

Masked Image Modeling (MIM) is a technique in self-supervised learning that focuses on acquiring detailed visual representations from unlabeled images by estimating the missing pixels in randomly masked sections. It has proven to be a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Khanh-Binh Nguyen , Chae Jung Park

Visual Object Tracking (VOT) is widely used in applications like autonomous driving to continuously track targets in videos. Existing methods can be roughly categorized into template matching and autoregressive methods, where the former…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Qianxiong Xu , Lanyun Zhu , Chenxi Liu , Guosheng Lin , Cheng Long , Ziyue Li , Rui Zhao

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the…

Robotics · Computer Science 2025-07-30 Fangqi Zhu , Hongtao Wu , Song Guo , Yuxiao Liu , Chilam Cheang , Tao Kong

With the growing need to effectively support workforce upskilling in the manufacturing sector, virtual reality is gaining popularity as a scalable training solution. However, most current systems are designed as static, step-by-step…

Human-Computer Interaction · Computer Science 2025-10-08 Bhavya Matam , Adamay Mann , Kachina Studer , Christian Gabbianelli , Sonia Castelo , John Liu , Claudio Silva , Dishita Turakhia

Occlusions caused by a robot's own body is a common problem for closed-loop control methods employed in eye-to-hand camera setups. We propose an optimization-based reactive controller that minimizes self-occlusions while achieving a desired…

In this contribution, we introduce the concept of Instance Performance Difference (IPD), a metric designed to measure the gap in performance that a robotics perception task experiences when working with real vs. synthetic pictures. By…

Robotics · Computer Science 2024-11-13 Bo-Hsun Chen , Dan Negrut

Identifying predictive world models for robots in novel environments from sparse online observations is essential for robot task planning and execution in novel environments. However, existing methods that leverage differentiable…

Robotics · Computer Science 2025-05-13 Yifan Zhu , Tianyi Xiang , Aaron Dollar , Zherong Pan

Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Kaicong Huang , Talha Azfar , Weisong Shi , Ruimin Ke

Video depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient and lacks robustness. An alternative is learning to enforce…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Yiran Wang , Min Shi , Jiaqi Li , Chaoyi Hong , Zihao Huang , Juewen Peng , Zhiguo Cao , Jianming Zhang , Ke Xian , Guosheng Lin

We present $multipanda\_ros2$, a novel open-source ROS2 architecture for multi-robot control of Franka Robotics robots. Leveraging ros2 control, this framework provides native ROS2 interfaces for controlling any number of robots from a…

Robotics · Computer Science 2026-02-03 Jon Škerlj , Seongjin Bien , Abdeldjallil Naceri , Sami Haddadin