English
Related papers

Related papers: StereoMV2D: A Sparse Temporal Stereo-Enhanced Fram…

200 papers

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Longlong Jing , Ruichi Yu , Henrik Kretzschmar , Kang Li , Charles R. Qi , Hang Zhao , Alper Ayvaci , Xu Chen , Dillon Cower , Yingwei Li , Yurong You , Han Deng , Congcong Li , Dragomir Anguelov

3D object detection at long range is crucial for ensuring the safety and efficiency of self driving vehicles, allowing them to accurately perceive and react to objects, obstacles, and potential hazards from a distance. But most current…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Ajinkya Khoche , Laura Pereira Sánchez , Nazre Batool , Sina Sharif Mansouri , Patric Jensfelt

Deep learning based 3D shape generation methods generally utilize latent features extracted from color images to encode the semantics of objects and guide the shape generation process. These color image semantics only implicitly encode 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Rakesh Shrestha , Zhiwen Fan , Qingkun Su , Zuozhuo Dai , Siyu Zhu , Ping Tan

Annotating 3D data remains a costly bottleneck for 3D object detection, motivating the development of weakly supervised annotation methods that rely on more accessible 2D box annotations. However, relying solely on 2D boxes introduces…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Saad Lahlali , Alexandre Fournier Montgieux , Nicolas Granger , Hervé Le Borgne , Quoc Cuong Pham

Multi-view cooperative perception and multimodal fusion are essential for reliable 3D spatiotemporal understanding in autonomous driving, especially under occlusions, limited viewpoints, and communication delays in V2X scenarios. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Zhenwei Yang , Yibo Ai , Weidong Zhang

Current multi-view 3D object detection methods often fail to detect objects in the overlap region properly, and the networks' understanding of the scene is often limited to that of a monocular detection network. Moreover, objects in the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Wonseok Roh , Gyusam Chang , Seokha Moon , Giljoo Nam , Chanyoung Kim , Younghyun Kim , Jinkyu Kim , Sangpil Kim

Recent video depth estimation methods achieve great performance by following the paradigm of image depth estimation, i.e., typically fine-tuning pre-trained video diffusion models with massive data. However, we argue that video depth…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Haodong Li , Chen Wang , Jiahui Lei , Kostas Daniilidis , Lingjie Liu

LiDAR-produced point clouds are the major source for most state-of-the-art 3D object detectors. Yet, small, distant, and incomplete objects with sparse or few points are often hard to detect. We present Sparse2Dense, a new framework to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Tianyu Wang , Xiaowei Hu , Zhengzhe Liu , Chi-Wing Fu

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Shuo Wang , Xinhai Zhao , Hai-Ming Xu , Zehui Chen , Dameng Yu , Jiahao Chang , Zhen Yang , Feng Zhao

Stereo is a prominent technique to infer dense depth maps from images, and deep learning further pushed forward the state-of-the-art, making end-to-end architectures unrivaled when enough data is available for training. However, deep…

Computer Vision and Pattern Recognition · Computer Science 2019-05-27 Matteo Poggi , Davide Pallotti , Fabio Tosi , Stefano Mattoccia

We present Multi-View Attentive Contextualization (MvACon), a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Xianpeng Liu , Ce Zheng , Ming Qian , Nan Xue , Chen Chen , Zhebin Zhang , Chen Li , Tianfu Wu

We propose Differentiable Stereopsis, a multi-view stereo approach that reconstructs shape and texture from few input views and noisy cameras. We pair traditional stereopsis and modern differentiable rendering to build an end-to-end model…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Shubham Goel , Georgia Gkioxari , Jitendra Malik

4D automotive radar is indispensable for autonomous driving due to its low cost and robustness, yet its point cloud sparsity challenges 3D object detection. Existing 4D radar-camera fusion methods focus on complex fusion strategies, trading…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Weiyi Xiong , Bing Zhu

Multi-view stereo reconstruction (MVS) in the wild requires to first estimate the camera parameters e.g. intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangulate…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Shuzhe Wang , Vincent Leroy , Yohann Cabon , Boris Chidlovskii , Jerome Revaud

Existing multi-view three-dimensional (3D) object detection approaches widely adopt large-scale pre-trained vision transformer (ViT)-based foundation models as backbones, being computationally complex. To address this problem, current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Danish Nazir , Antoine Hanna-Asaad , Lucas Görnhardt , Jan Piewek , Thorsten Bagdonat , Tim Fingscheidt

LiDAR-based 3D object detection presents significant challenges due to the inherent sparsity of LiDAR points. A common solution involves long-term temporal LiDAR data to densify the inputs. However, efficiently leveraging spatial-temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Chaoqun Wang , Xiaobin Hong , Wenzhong Li , Ruimao Zhang

Autonomous driving perception tasks rely heavily on cameras as the primary sensor for Object Detection, Semantic Segmentation, Instance Segmentation, and Object Tracking. However, RGB images captured by cameras lack depth information, which…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Marcelo Eduardo Pederiva , José Mario De Martino , Alessandro Zimmer

Recent advances in robot imitation learning have yielded powerful visuomotor policies capable of manipulating a wide variety of objects directly from monocular visual inputs. However, monocular observations inherently lack reliable depth…

Robotics · Computer Science 2026-05-12 Evans Han , Yunfan Jiang , Yingke Wang , Haoyue Xiao , Huang Huang , Jianwen Xie , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

Efficient and accurate 3D reconstruction is crucial for various applications, including augmented and virtual reality, medical imaging, and cinematic special effects. While traditional Multi-View Stereo (MVS) systems have been fundamental…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Umair Haroon , Ahmad AlMughrabi , Ricardo Marques , Petia Radeva

The fusion of LiDAR and camera sensors has demonstrated significant effectiveness in achieving accurate detection for short-range tasks in autonomous driving. However, this fusion approach could face challenges when dealing with long-range…

Image and Video Processing · Electrical Eng. & Systems 2025-03-27 Tanmoy Dam , Sanjay Bhargav Dharavath , Sameer Alam , Nimrod Lilith , Aniruddha Maiti , Supriyo Chakraborty , Mir Feroskhan