English
Related papers

Related papers: PosePilot: Steering Camera Pose for Generative Wor…

200 papers

Tracking a point through a video can be a challenging task due to uncertainty arising from visual obfuscations, such as appearance changes and occlusions. Although current state-of-the-art discriminative models excel in regressing long-term…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Mattie Tesfaldet , Adam W. Harley , Konstantinos G. Derpanis , Derek Nowrouzezahrai , Christopher Pal

3D pose estimation has recently gained substantial interests in computer vision domain. Existing 3D pose estimation methods have a strong reliance on large size well-annotated 3D pose datasets, and they suffer poor model generalization on…

Computer Vision and Pattern Recognition · Computer Science 2022-07-11 Shannan Guan , Haiyan Lu , Linchao Zhu , Gengfa Fang

Depth estimation is a critical technology in autonomous driving, and multi-camera systems are often used to achieve a 360$^\circ$ perception. These 360$^\circ$ camera sets often have limited or low-quality overlap regions, making multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Jialei Xu , Wei Yin , Dong Gong , Junjun Jiang , Xianming Liu

This paper presents DriVerse, a generative model for simulating navigation-driven driving scenes from a single image and a future trajectory. Previous autonomous driving world models either directly feed the trajectory or discrete control…

Robotics · Computer Science 2026-04-28 Xiaofan Li , Chenming Wu , Zhao Yang , Zhihao Xu , Dingkang Liang , Yumeng Zhang , Ji Wan , Jun Wang

Over the last two decades, deep learning has transformed the field of computer vision. Deep convolutional networks were successfully applied to learn different vision tasks such as image classification, image segmentation, object detection…

Computer Vision and Pattern Recognition · Computer Science 2019-07-17 Yoli Shavit , Ron Ferens

World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Yuntao Chen , Yuqi Wang , Zhaoxiang Zhang

Robust trajectory planning under camera viewpoint changes is important for scalable end-to-end autonomous driving. However, existing models often depend heavily on the camera viewpoints seen during training. We investigate an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Hiroki Hashimoto , Hiromichi Goto , Hiroyuki Sugai , Hiroshi Kera , Kazuhiko Kawamoto

Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Zihan Zhou , Luxi Chen , Jingzhi Zhou , Yuhao Wan , Min Zhao , Baoyu Fan , Chongxuan Li

Learning world models can teach an agent how the world works in an unsupervised manner. Even though it can be viewed as a special case of sequence modeling, progress for scaling world models on robotic applications such as autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Lunjun Zhang , Yuwen Xiong , Ze Yang , Sergio Casas , Rui Hu , Raquel Urtasun

This paper introduces PoseLess, a novel framework for robot hand control that eliminates the need for explicit pose estimation by directly mapping 2D images to joint angles using projected representations. Our approach leverages synthetic…

Robotics · Computer Science 2025-03-12 Alan Dao , Dinh Bach Vu , Tuan Le Duc Anh , Bui Quang Huy

Pose estimation and map building are central ingredients of autonomous robots and typically rely on the registration of sensor data. In this paper, we investigate a new metric for registering images that builds upon on the idea of the…

Computer Vision and Pattern Recognition · Computer Science 2020-04-09 Jan Quenzel , Radu Alexandru Rosu , Thomas Läbe , Cyrill Stachniss , Sven Behnke

Given sparse views of a 3D object, estimating their camera poses is a long-standing and intractable problem. Toward this goal, we consider harnessing the pre-trained diffusion model of novel views conditioned on viewpoints (Zero-1-to-3). We…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Weihao Cheng , Yan-Pei Cao , Ying Shan

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Sen Wang , Jingyi Tian , Le Wang , Zhimin Liao , Jiayi Li , Huaiyi Dong , Kun Xia , Sanping Zhou , Wei Tang , Hua Gang

Unsupervised monocular depth estimation frameworks have shown promising performance in autonomous driving. However, existing solutions primarily rely on a simple convolutional neural network for ego-motion recovery, which struggles to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Yi Feng , Zizhan Guo , Qijun Chen , Rui Fan

As 360{\deg} cameras become prevalent in many autonomous systems (e.g., self-driving cars and drones), efficient 360{\deg} perception becomes more and more important. We propose a novel self-supervised learning approach for predicting the…

Computer Vision and Pattern Recognition · Computer Science 2018-11-14 Fu-En Wang , Hou-Ning Hu , Hsien-Tzu Cheng , Juan-Ting Lin , Shang-Ta Yang , Meng-Li Shih , Hung-Kuo Chu , Min Sun

Real-time robotic grasping, supporting a subsequent precise object-in-hand operation task, is a priority target towards highly advanced autonomous systems. However, such an algorithm which can perform sufficiently-accurate grasping with…

Computer Vision and Pattern Recognition · Computer Science 2021-11-12 Tuan-Tang Le , Trung-Son Le , Yu-Ru Chen , Joel Vidal , Chyi-Yeu Lin

Most successful approaches to estimate the 6D pose of an object typically train a neural network by supervising the learning with annotated poses in real world images. These annotations are generally expensive to obtain and a common…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Juil Sock , Guillermo Garcia-Hernando , Anil Armagan , Tae-Kyun Kim

Pose-guided video generation has become a powerful tool in creative industries, exemplified by frameworks like Animate Anyone. However, conditioning generation on specific poses introduces serious risks, such as impersonation, privacy…

Cryptography and Security · Computer Science 2025-08-05 Kongxin Wang , Jie Zhang , Peigui Qi , Kunsheng Tang , Tianwei Zhang , Wenbo Zhou

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

Artificial Intelligence · Computer Science 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao

Real-world robotics applications demand object pose estimation methods that work reliably across a variety of scenarios. Modern learning-based approaches require large labeled datasets and tend to perform poorly outside the training domain.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-15 Jingnan Shi , Rajat Talak , Dominic Maggio , Luca Carlone
‹ Prev 1 4 5 6 7 8 10 Next ›