English
Related papers

Related papers: Geometric-aware Pretraining for Vision-centric 3D …

200 papers

Multi-object tracking (MOT) in monocular videos is fundamentally challenged by occlusions and depth ambiguity, issues that conventional tracking-by-detection (TBD) methods struggle to resolve owing to a lack of geometric awareness. To…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xudong Han , Pengcheng Fang , Yueying Tian , Jianhui Yu , Xiaohao Cai , Daniel Roggen , Philip Birch

The recent surge in popularity of deep generative models for 3D objects has highlighted the need for more efficient training methods, particularly given the difficulties associated with training with conventional 3D representations, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Adam Kania , Artur Kasymov , Jakub Kościukiewicz , Artur Górak , Marcin Mazur , Maciej Zięba , Przemysław Spurek

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen

4D radar has received significant attention in autonomous driving thanks to its robustness under adverse weathers. Due to the sparse points and noisy measurements of the 4D radar, most of the research finish the 3D object detection task by…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Hanzhi Zhong , Zhiyu Xiang , Ruoyu Xu , Jingyun Fu , Peng Xu , Shaohong Wang , Zhihao Yang , Tianyu Pu , Eryun Liu

We present recurrent geometry-aware neural networks that integrate visual information across multiple views of a scene into 3D latent feature tensors, while maintaining an one-to-one mapping between 3D physical locations in the world scene…

Computer Vision and Pattern Recognition · Computer Science 2018-11-15 Ricson Cheng , Ziyan Wang , Katerina Fragkiadaki

Recently, 3D object detection has attracted significant attention and achieved continuous improvement in real road scenarios. The environmental information is collected from a single sensor or multi-sensor fusion to detect interested…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Hongwei Liu , Jian Yang , Jianfeng Zhang , Dongheng Shao , Jielong Guo , Shaobo Li , Xuan Tang , Xian Wei

Learning powerful representations in bird's-eye-view (BEV) for perception tasks is trending and drawing extensive attention both from industry and academia. Conventional approaches for most autonomous driving algorithms perform detection,…

Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning for various downstream tasks, such as semantic segmentation,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yiren Lu , Xin Ye , Burhaneddin Yaman , Jingru Luo , Zhexiao Xiong , Liu Ren , Yu Yin

The pretraining-finetuning paradigm has powered major advances in domains such as natural language processing and computer vision, with representative examples including masked language modeling and next-token prediction. In molecular…

Machine Learning · Computer Science 2025-10-21 Shaoheng Yan , Zian Li , Muhan Zhang

Depth estimation is a cornerstone of perception in autonomous driving and robotic systems. The considerable cost and relatively sparse data acquisition of LiDAR systems have led to the exploration of cost-effective alternatives, notably,…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Yucheng Mao , Ruowen Zhao , Tianbao Zhang , Hang Zhao

3D object-level mapping is a fundamental problem in robotics, which is especially challenging when object CAD models are unavailable during inference. In this work, we propose a framework that can reconstruct high-quality object-level maps…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Ziwei Liao , Jun Yang , Jingxing Qian , Angela P. Schoellig , Steven L. Waslander

Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while elevating novel view quality. Due to the surround-view with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Junhong Lin , Kangli Wang , Shunzhou Wang , Songlin Fan , Ge Li , Wei Gao

Vision-based bird's-eye-view (BEV) 3D object detection has advanced significantly in autonomous driving by offering cost-effectiveness and rich contextual information. However, existing methods often construct BEV representations by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jicheng Yuan , Manh Nguyen Duc , Qian Liu , Manfred Hauswirth , Danh Le Phuoc

Unaligned Scene Change Detection aims to detect scene changes between image pairs captured at different times without assuming viewpoint alignment. To handle viewpoint variations, current methods rely solely on 2D visual cues to establish…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Ziling Liu , Ziwei Chen , Mingqi Gao , Jinyu Yang , Feng Zheng

In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer designs, which struggle to resolve geometric information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Ziwei Liao , Jialiang Zhu , Chunyu Wang , Han Hu , Steven L. Waslander

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Shuo Wang , Xinhai Zhao , Hai-Ming Xu , Zehui Chen , Dameng Yu , Jiahao Chang , Zhen Yang , Feng Zhao

Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yiyao Ma , Kai Chen , Zhongxiang Zhou , Zhuheng Song , Dongsheng Xie , Zelong Tan , Rong Xiong , Qi Dou

Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 John J. Han , Adam Schmidt , Muhammad Abdullah Jamal , Chinedu Nwoye , Anita Rau , Jie Ying Wu , Omid Mohareri

Accurate 3D lane detection from monocular images presents significant challenges due to depth ambiguity and imperfect ground modeling. Previous attempts to model the ground have often used a planar ground assumption with limited degrees of…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Chaesong Park , Eunbin Seo , Jongwoo Lim

This paper addresses key challenges in object-centric representation learning of video. While existing approaches struggle with complex scenes, we propose a novel weakly-supervised framework that emphasises geometric understanding and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Phúc H. Le Khac , Graham Healy , Alan F. Smeaton
‹ Prev 1 4 5 6 7 8 10 Next ›