English
Related papers

Related papers: OmniD: Generalizable Robot Manipulation Policy via…

200 papers

Diffusion strategies have advanced visual motor control by progressively denoising high-dimensional action sequences, providing a promising method for robot manipulation. However, as task complexity increases, the success rate of existing…

Robotics · Computer Science 2026-01-21 Weize Xie , Yi Ding , Ying He , Leilei Wang , Binwen Bai , Zheyi Zhao , Chenyang Wang , F. Richard Yu

End-to-end visuomotor policies trained using behavior cloning have shown a remarkable ability to generate complex, multi-modal low-level robot behaviors. However, at deployment time, these policies still struggle to act reliably when faced…

Robotics · Computer Science 2025-06-17 Pranay Gupta , Henny Admoni , Andrea Bajcsy

Imitation learning provides an efficient way to teach robots dexterous skills; however, learning complex skills robustly and generalizablely usually consumes large amounts of human demonstrations. To tackle this challenging problem, we…

Robotics · Computer Science 2024-09-30 Yanjie Ze , Gu Zhang , Kangning Zhang , Chenyuan Hu , Muhan Wang , Huazhe Xu

Does multi-view demonstration truly improve robot manipulation, or merely enhance cross-view robustness? We present a systematic study quantifying the performance gains, scaling behavior, and underlying mechanisms of multi-view data for…

Infrared and visible image fusion (IVIF) is increasingly applied in critical fields such as video surveillance and autonomous driving systems. Significant progress has been made in deep learning-based fusion methods. However, these models…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Yukai Shi , Cidan Shi , Zhipeng Weng , Yin Tian , Xiaoyu Xian , Liang Lin

Visual bird's eye view (BEV) perception, due to its excellent perceptual capabilities, is progressively replacing costly LiDAR-based perception systems, especially in the realm of urban intelligent driving. However, this type of perception…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Lei He , Qiaoyi Wang , Honglin Sun , Qing Xu , Bolin Gao , Shengbo Eben Li , Jianqiang Wang , Keqiang Li

In the landscape of autonomous driving, Bird's-Eye-View (BEV) representation has recently garnered substantial academic attention, serving as a transformative framework for the fusion of multi-modal sensor inputs. This BEV paradigm…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Yuxin Li , Yiheng Li , Xulei Yang , Mengying Yu , Zihang Huang , Xiaojun Wu , Chai Kiat Yeo

Diffusion policies excel at visuomotor control but often fail catastrophically under severe out-of-distribution (OOD) disturbances, such as unexpected object displacements or visual corruptions. To address this vulnerability, we introduce…

Robotics · Computer Science 2026-03-24 Ziou Hu , Xiangtong Yao , Yuan Meng , Zhenshan Bing , Alois Knoll

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive…

Computation and Language · Computer Science 2026-04-02 Jaeik Kim , Woojin Kim , Jihwan Hong , Yejoon Lee , Sieun Hyeon , Mintaek Lim , Yunseok Han , Dogeun Kim , Hoeun Lee , Hyunggeun Kim , Jaeyoung Do

Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, current methods underutilize their distribution modeling…

This paper introduces InverseMatrixVT3D, an efficient method for transforming multi-view image features into 3D feature volumes for 3D semantic occupancy prediction. Existing methods for constructing 3D volumes often rely on depth…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Zhenxing Ming , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Haoteng Li , Zhao Yang , Zezhong Qian , Gongpeng Zhao , Yuqi Huang , Jun Yu , Huazheng Zhou , Longjun Liu

Bird's-eye-view (BEV) semantic segmentation is becoming crucial in autonomous driving systems. It realizes ego-vehicle surrounding environment perception by projecting 2D multi-view images into 3D world space. Recently, BEV segmentation has…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Jian Sun , Yuqi Dai , Chi-Man Vong , Qing Xu , Shengbo Eben Li , Jianqiang Wang , Lei He , Keqiang Li

Learning robust visuomotor policies that generalize across diverse objects and interaction dynamics remains a central challenge in robotic manipulation. Most existing approaches rely on direct observation-to-action mappings or compress…

Robotics · Computer Science 2025-09-24 Sangjun Noh , Dongwoo Nam , Kangmin Kim , Geonhyup Lee , Yeonguk Yu , Raeyoung Kang , Kyoobin Lee

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Rohit Mohan , Daniele Cattaneo , Florian Drews , Abhinav Valada

Bird's-Eye-View (BEV) perception has become a vital component of autonomous driving systems due to its ability to integrate multiple sensor inputs into a unified representation, enhancing performance in various downstream tasks. However,…

Robotics · Computer Science 2024-10-10 Yuxin Li , Yiheng Li , Xulei Yang , Mengying Yu , Zihang Huang , Xiaojun Wu , Chai Kiat Yeo

Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint changes, and cross-embodiment transfer, yet they are…

Robotics · Computer Science 2026-01-28 Ruiyu Wang , Zheyu Zhuang , Danica Kragic , Florian T. Pokorny

Prior research on out-of-distribution detection (OoDD) has primarily focused on single-modality models. Recently, with the advent of large-scale pretrained vision-language models such as CLIP, OoDD methods utilizing such multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Jeonghyeon Kim , Sangheum Hwang

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained…

LiDAR scene generation is increasingly important for scalable simulation and synthetic data creation, especially under diverse sensing conditions that are costly to capture at scale. Typically, diffusion-based LiDAR generators are developed…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Youquan Liu , Weidong Yang , Ao Liang , Xiang Xu , Lingdong Kong , Yang Wu , Dekai Zhu , Xin Li , Runnan Chen , Ben Fei , Tongliang Liu , Wanli Ouyang