English
Related papers

Related papers: OmniD: Generalizable Robot Manipulation Policy via…

200 papers

High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. However, current policies often hit a "generality barrier": as motion libraries scale in diversity, tracking fidelity inevitably…

Scooping items with tools such as spoons and ladles is common in daily life, ranging from assistive feeding to retrieving items from environmental disaster sites. However, developing a general and autonomous robotic scooping policy is…

Robotics · Computer Science 2025-10-14 Kuanning Wang , Yongchong Gu , Yuqian Fu , Zeyu Shangguan , Sicheng He , Xiangyang Xue , Yanwei Fu , Daniel Seita

Inspired by the success of transfer learning in computer vision, roboticists have investigated visual pre-training as a means to improve the learning efficiency and generalization ability of policies learned from pixels. To that end, past…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Kaylee Burns , Zach Witzel , Jubayer Ibn Hamid , Tianhe Yu , Chelsea Finn , Karol Hausman

The generalization ability of visuomotor policy is crucial, as a good policy should be deployable across diverse scenarios. Some methods can collect large amounts of trajectory augmentation data to train more generalizable imitation…

Robotics · Computer Science 2025-11-14 Hanwen Wang

Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generation task, lacking explicit 3D modeling. However, we argue…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Zeming Chen , Hang Zhao

Recognizing out-of-distribution (OOD) samples is critical for machine learning systems deployed in the open world. The vast majority of OOD detection methods are driven by a single modality (e.g., either vision or language), leaving the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Yifei Ming , Ziyang Cai , Jiuxiang Gu , Yiyou Sun , Wei Li , Yixuan Li

In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Zerui Chen , Shizhe Chen , Etienne Arlaud , Ivan Laptev , Cordelia Schmid

The bird's-eye-view (BEV) representation allows robust learning of multiple tasks for autonomous driving including road layout estimation and 3D object detection. However, contemporary methods for unified road layout estimation and 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Curie Kim , Ue-Hwan Kim

Learned visuomotor policies are capable of performing increasingly complex manipulation tasks. However, most of these policies are trained on data collected from limited robot positions and camera viewpoints. This leads to poor…

Robotics · Computer Science 2025-09-29 Jingyun Yang , Isabella Huang , Brandon Vu , Max Bajracharya , Rika Antonova , Jeannette Bohg

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. Existing models, however, are typically restricted to limited state modalities, short video sequences, imprecise…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Bohan Li , Zhuang Ma , Dalong Du , Baorui Peng , Zhujin Liang , Zhenqiang Liu , Chao Ma , Yueming Jin , Hao Zhao , Wenjun Zeng , Xin Jin

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it…

Robotics · Computer Science 2025-07-01 Zihan Lan , Weixin Mao , Haosheng Li , Le Wang , Tiancai Wang , Haoqiang Fan , Osamu Yoshie

Recent vision-only perception models for autonomous driving achieved promising results by encoding multi-view image features into Bird's-Eye-View (BEV) space. A critical step and the main bottleneck of these methods is transforming image…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Jiayu Yang , Enze Xie , Miaomiao Liu , Jose M. Alvarez

Generating consistent multiple views for 3D reconstruction tasks is still a challenge to existing image-to-3D diffusion models. Generally, incorporating 3D representations into diffusion model decrease the model's speed as well as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Emmanuelle Bourigault , Pauline Bourigault

Camera-based end-to-end driving neural networks bring the promise of a low-cost system that maps camera images to driving control commands. These networks are appealing because they replace laborious hand engineered building blocks but…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Abdelhak Loukkal , Yves Grandvalet , Tom Drummond , You Li

Learning visual representations from observing actions to benefit robot visuo-motor policy generation is a promising direction that closely resembles human cognitive function and perception. Motivated by this, and further inspired by…

LiDAR and camera are two essential sensors for 3D object detection in autonomous driving. LiDAR provides accurate and reliable 3D geometry information while the camera provides rich texture with color. Despite the increasing popularity of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Qi Jiang , Hao Sun , Xi Zhang

Feed-forward 3D Gaussian splatting (3DGS) models have gained significant popularity due to their ability to generate scenes immediately without needing per-scene optimization. Although omnidirectional images are becoming more popular since…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Suyoung Lee , Jaeyoung Chung , Kihoon Kim , Jaeyoo Huh , Gunhee Lee , Minsoo Lee , Kyoung Mu Lee

Autonomous driving requires an understanding of the static environment from sensor data. Learned Bird's-Eye View (BEV) encoders are commonly used to fuse multiple inputs, and a vector decoder predicts a vectorized map representation from…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Thomas Monninger , Zihan Zhang , Zhipeng Mo , Md Zafar Anwar , Steffen Staab , Sihao Ding

Embodied visual tracking is to follow a target object in dynamic 3D environments using an agent's egocentric vision. This is a vital and challenging skill for embodied agents. However, existing methods suffer from inefficient training and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Fangwei Zhong , Kui Wu , Hai Ci , Churan Wang , Hao Chen

Scaling up robot learning requires large and diverse datasets, and how to efficiently reuse collected data and transfer policies to new embodiments remains an open question. Emerging research such as the Open-X Embodiment (OXE) project has…