中文
相关论文

相关论文: Can These Views Be One Scene? Evaluating Multiview…

200 篇论文

Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high-precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3D foundation models…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Kerui Ren , Guanghao Li , Changjian Jiang , Yingxiang Xu , Tao Lu , Linning Xu , Junting Dong , Jiangmiao Pang , Mulin Yu , Bo Dai

We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (using vision language models). Recently, MVGs have been the…

图形学 · 计算机科学 2025-07-02 Xianghui Xie , Chuhang Zou , Meher Gitika Karumuri , Jan Eric Lenssen , Gerard Pons-Moll

Recent advances in data-driven geometric multi-view 3D reconstruction foundation models (e.g., DUSt3R) have shown remarkable performance across various 3D vision tasks, facilitated by the release of large-scale, high-quality 3D datasets.…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Wenyu Li , Sidun Liu , Peng Qiao , Yong Dou

Multi-View Stereo (MVS) is a core task in 3D computer vision. With the surge of novel deep learning methods, learned MVS has surpassed the accuracy of classical approaches, but still relies on building a memory intensive dense cost volume.…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Radu Alexandru Rosu , Sven Behnke

Given a single image of a 3D object, this paper proposes a novel method (named ConsistNet) that is able to generate multiple images of the same object, as if seen they are captured from different viewpoints, while the 3D (multi-view)…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Jiayu Yang , Ziang Cheng , Yunfei Duan , Pan Ji , Hongdong Li

Traditional multi-view stereo (MVS) methods primarily depend on photometric and geometric consistency constraints. In contrast, modern learning-based algorithms often rely on the plane sweep algorithm to infer 3D geometry, applying explicit…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Vibhas Vats , Md. Alimoor Reza , David Crandall , Soon-heung Jung

The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions reflect stable multimodal processing. In this work, we argue that this assumption is…

Robots often rely on RGB images for tasks like manipulation and navigation. However, reliable interaction typically requires a 3D scene representation that is metric-scaled and aligned with the robot reference frame. This depends on…

机器人学 · 计算机科学 2025-09-11 Davide Allegro , Matteo Terreran , Stefano Ghidoni

Humans are able to accurately reason in 3D by gathering multi-view observations of the surrounding world. Inspired by this insight, we introduce a new large-scale benchmark for 3D multi-view visual question answering (3DMV-VQA). This…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Yining Hong , Chunru Lin , Yilun Du , Zhenfang Chen , Joshua B. Tenenbaum , Chuang Gan

Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applications such as immersive interaction and content creation. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Xiaoxuan Ma , Jiashun Wang , Nicolas Ugrinovic , Yehonathan Litman , Kris Kitani

We introduce RealX3D, a real-capture benchmark for multi-view visual restoration and 3D reconstruction under diverse physical degradations. RealX3D groups corruptions into four families, including illumination, scattering, occlusion, and…

Novel view synthesis is an important problem with many applications, including AR/VR, gaming, and robotic simulations. With the recent rapid development of Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) methods, it is…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Jonas Kulhanek , Torsten Sattler

3D shape completion is important to enable machines to perceive the complete geometry of objects from partial observations. To address this problem, view-based methods have been presented. These methods represent shapes as multiple depth…

计算机视觉与模式识别 · 计算机科学 2019-12-02 Tao Hu , Zhizhong Han , Matthias Zwicker

A longstanding question in computer vision concerns the representation of 3D shapes for recognition: should 3D shapes be represented with descriptors operating on their native 3D formats, such as voxel grid or polygon mesh, or can they be…

计算机视觉与模式识别 · 计算机科学 2015-09-29 Hang Su , Subhransu Maji , Evangelos Kalogerakis , Erik Learned-Miller

Neural Radiance Fields (NeRF) has demonstrated remarkable 3D reconstruction capabilities with dense view images. However, its performance significantly deteriorates under sparse view settings. We observe that learning the 3D consistency of…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Shoukang Hu , Kaichen Zhou , Kaiyu Li , Longhui Yu , Lanqing Hong , Tianyang Hu , Zhenguo Li , Gim Hee Lee , Ziwei Liu

Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet they often underperform strong per-frame methods on…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Guangkai Xu , Hua Geng , Huanyi Zheng , Songyi Yin , Yanlong Sun , Hao Chen , Chunhua Shen

To what extent are two images picturing the same 3D surfaces? Even when this is a known scene, the answer typically requires an expensive search across scale space, with matching and geometric verification of large sets of local features.…

计算机视觉与模式识别 · 计算机科学 2020-08-14 Anita Rau , Guillermo Garcia-Hernando , Danail Stoyanov , Gabriel J. Brostow , Daniyar Turmukhambetov

Remarkable progress has been made in self-supervised monocular depth estimation (SS-MDE) by exploring cross-view consistency, e.g., photometric consistency and 3D point cloud consistency. However, they are very vulnerable to illumination…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Haimei Zhao , Jing Zhang , Zhuo Chen , Bo Yuan , Dacheng Tao

Multi-modal 3D object detection with bird's eye view (BEV) has achieved desired advances on benchmarks. Nonetheless, the accuracy may drop significantly in the real world due to data corruption such as sensor configurations for LiDAR and…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Rui Ding , Zhaonian Kuang , Yuzhe Ji , Meng Yang , Xinhu Zheng , Gang Hua

Vision-Language Models (VLMs) have achieved impressive performance across a wide range of multimodal tasks, yet they often exhibit inconsistent behavior when faced with semantically equivalent inputs, undermining their reliability and…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Shih-Han Chou , Shivam Chandhok , James J. Little , Leonid Sigal