English
Related papers

Related papers: Unlocking the Power of Critical Factors for 3D Vis…

200 papers

Deep Convolutional Neural Networks (DCNNs) and their variants have been widely used in large scale face recognition(FR) recently. Existing methods have achieved good performance on many FR benchmarks. However, most of them suffer from two…

Computer Vision and Pattern Recognition · Computer Science 2021-06-28 Jing Xu , Tszhang Guo , Yong Xu , Zenglin Xu , Kun Bai

Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches operate purely in the 2D pixel space and lack explicit…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yunpeng Bai , Shaoheng Fang , Chaohui Yu , Fan Wang , Qixing Huang

Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale prediction, enabling high-quality image generation. However, the VAR paradigm suffers from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Senmao Li , Kai Wang , Salman Khan , Fahad Shahbaz Khan , Jian Yang , Yaxing Wang

Representing 3D scenes from multiview images is a core challenge in computer vision and graphics, which requires both precise rendering and accurate reconstruction. Recently, 3D Gaussian Splatting (3DGS) has garnered significant attention…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 You Shen , Zhipeng Zhang , Xinyang Li , Yansong Qu , Yu Lin , Shengchuan Zhang , Liujuan Cao

Cross-view localization aims to estimate the 3-DoF pose of a ground-view image by aligning it with aerial or satellite imagery. Existing methods typically address this task through direct regression or feature alignment in a shared…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Panwang Xia , Qiong Wu , Lei Yu , Yi Liu , Mingtao Xiong , Xudong Lu , Yi Liu , Haoyu Guo , Yongxiang Yao , Junjian Zhang , Xiangyuan Cai , Hongwei Hu , Zhi Zheng , Yongjun Zhang , Yi Wan

In this work, we address the problem of refining the geometry of local image features from multiple views without known scene or camera geometry. Current approaches to local feature detection are inherently limited in their keypoint…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Mihai Dusmanu , Johannes L. Schönberger , Marc Pollefeys

The self-localization capability is a crucial component for Unmanned Ground Vehicles (UGV) in farming applications. Approaches based solely on visual cues or on low-cost GPS are easily prone to fail in such scenarios. In this paper, we…

Robotics · Computer Science 2018-09-12 Marco Imperoli , Ciro Potena , Daniele Nardi , Giorgio Grisetti , Alberto Pretto

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-world 3D assets with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xinyue Liang , Zhinyuan Ma , Lingchen Sun , Yanjun Guo , Lei Zhang

We address the problem of recovering the 3D geometry of a human face from a set of facial images in multiple views. While recent studies have shown impressive progress in 3D Morphable Model (3DMM) based facial reconstruction, the settings…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Fanzi Wu , Linchao Bao , Yajing Chen , Yonggen Ling , Yibing Song , Songnan Li , King Ngi Ngan , Wei Liu

Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometrically meaningful content. Existing approaches typically learn a mapping from camera…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Chen Hou , Christian Rupprecht

Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Shuo Wang , Xinhai Zhao , Hai-Ming Xu , Zehui Chen , Dameng Yu , Jiahao Chang , Zhen Yang , Feng Zhao

The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Qitao Zhao , Ce Zheng , Mengyuan Liu , Chen Chen

Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal large language models (MLLMs) often struggle to reason about 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Om Khangaonkar , Hadi J. Rad , Hamed Pirsiavash

In this paper, we propose a new deep architecture for fusing camera and LiDAR sensors for 3D object detection. Because the camera and LiDAR sensor signals have different characteristics and distributions, fusing these two modalities is…

Computer Vision and Pattern Recognition · Computer Science 2020-12-10 Jin Hyeok Yoo , Yecheol Kim , Jisong Kim , Jun Won Choi

Recent advances in feature learning have shown that self-supervised vision foundation models can capture semantic correspondences but often lack awareness of underlying 3D geometry. GECO addresses this gap by producing geometrically…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Regine Hartwig , Dominik Muhle , Riccardo Marin , Daniel Cremers

Surround depth estimation provides a cost-effective alternative to LiDAR for 3D perception in autonomous driving. While recent self-supervised methods explore multi-camera settings to improve scale awareness and scene coverage, they are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Weimin Liu , Jiyuan Qiu , Wenjun Wang , Joshua H. Meng

Multi-modal large language models (MLLMs), such as GPT-4o, excel at integrating text and visual data but face systematic challenges when interpreting ambiguous or incomplete visual stimuli. This study leverages statistical modeling to…

Machine Learning · Computer Science 2024-12-09 Ching-Yi Wang

One of the greatest challenges in the design of a real-time perception system for autonomous driving vehicles and drones is the conflicting requirement of safety (high prediction accuracy) and efficiency. Traditional approaches use a single…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Ziyao Tang , Yongxi Lu , Tara Javidi

Computer Vision practitioners must thoroughly understand their model's performance, but conditional evaluation is complex and error-prone. In biometric verification, model performance over continuous covariates---real-number attributes of…

Machine Learning · Computer Science 2020-09-22 Mel McCurrie , Hamish Nicholson , Walter J. Scheirer , Samuel Anthony

Accurate camera viewpoint estimation under sparse-view conditions remains challenging, particularly in two-view scenarios. Recent approaches leverage diffusion models such as Zero123 to synthesize novel views conditioned on relative…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yan-Ting Chen , Hao-Wei Chen , Tsu-Ching Hsiao , Chun-Yi Lee