中文
相关论文

相关论文: H3R: Hybrid Multi-view Correspondence for Generali…

200 篇论文

In recent years, implicit online dense mapping methods have achieved high-quality reconstruction results, showcasing great potential in robotics, AR/VR, and digital twins applications. However, existing methods struggle with slow texture…

机器人学 · 计算机科学 2024-07-17 Chenxing Jiang , Yiming Luo , Boyu Zhou , Shaojie Shen

Recent advancements in neural visual geometry, including transformer-based models such as VGGT and Pi3, have achieved impressive accuracy on 3D reconstruction tasks. However, their reliance on full attention makes them fundamentally limited…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Leo Kaixuan Cheng , Abdus Shaikh , Ruofan Liang , Zhijie Wu , Yushi Guan , Nandita Vijaykumar

This paper investigates a 2D to 3D image translation method with a straightforward technique, enabling correlated 2D X-ray to 3D CT-like reconstruction. We observe that existing approaches, which integrate information across multiple 2D…

计算机视觉与模式识别 · 计算机科学 2024-06-27 Abril Corona-Figueroa , Hubert P. H. Shum , Chris G. Willcocks

Modern Earth Observation systems provide sensing data at different temporal and spatial resolutions. Among optical sensors, today the Sentinel-2 program supplies high-resolution temporal (every 5 days) and high spatial resolution (10m)…

计算机视觉与模式识别 · 计算机科学 2018-03-07 P. Benedetti , D. Ienco , R. Gaetano , K. Osé , R. Pensa , S. Dupuy

With the daily influx of 3D data on the internet, text-3D retrieval has gained increasing attention. However, current methods face two major challenges: Hierarchy Representation Collapse (HRC) and Redundancy-Induced Saliency Dilution…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Wenrui Li , Yidan Lu , Yeyu Chai , Rui Zhao , Hengyu Man , Xiaopeng Fan

Accurate Digital Surface Model (DSM) reconstruction from satellite imagery is critical for applications such as disaster response, urban planning, and large-scale geographic mapping. Existing approaches face a fundamental trade-off:…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Qiaoyi Yang , Chaoyi Zhou , Xi Liu , Run Wang , Minghui Xu , Mert D. Pesé , Feng Luo , Yuhao Xu , Zhi-Qi Cheng , Qiushi Chen , Hairong Qi , Siyu Huang

Visual correspondence across image-to-image (2D-2D), image-to-point cloud (2D-3D), and point cloud-to-point cloud (3D-3D) geometric matching forms the foundation for numerous 3D vision tasks. Despite sharing a similar problem structure,…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Prajnan Goswami , Tianye Ding , Feng Liu , Huaizu Jiang

We introduce DiHuR, a novel Diffusion-guided model for generalizable Human 3D Reconstruction and view synthesis from sparse, minimally overlapping images. While existing generalizable human radiance fields excel at novel view synthesis,…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Jinnan Chen , Chen Li , Gim Hee Lee

Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high-precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3D foundation models…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Kerui Ren , Guanghao Li , Changjian Jiang , Yingxiang Xu , Tao Lu , Linning Xu , Junting Dong , Jiangmiao Pang , Mulin Yu , Bo Dai

With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However, the excessive reliance on multi-view geometric annotations,…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Youyu Chen , Junjun Jiang , Yueru Luo , Kui Jiang , Xianming Liu , Xu Yan , Dave Zhenyu Chen

Multi-view transformers such as DUSt3R are revolutionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the inner mechanisms of multi-view transformers are unclear.…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Michal Stary , Julien Gaubil , Ayush Tewari , Vincent Sitzmann

In this paper, we introduce Splatt3R, a pose-free, feed-forward method for in-the-wild 3D reconstruction and novel view synthesis from stereo pairs. Given uncalibrated natural images, Splatt3R can predict 3D Gaussian Splats without…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Brandon Smart , Chuanxia Zheng , Iro Laina , Victor Adrian Prisacariu

We introduce G-CUT3R, a novel feed-forward approach for guided 3D scene reconstruction that enhances the CUT3R model by integrating prior information. Unlike existing feed-forward methods that rely solely on input images, our method…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ramil Khafizov , Artem Komarichev , Ruslan Rakhimov , Peter Wonka , Evgeny Burnaev

Traditionally, creating photo-realistic 3D head avatars requires a studio-level multi-view capture setup and expensive optimization during test-time, limiting the use of digital human doubles to the VFX industry or offline renderings. To…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Tobias Kirschstein , Javier Romero , Artem Sevastopolsky , Matthias Nießner , Shunsuke Saito

Scene reconstruction has emerged as a central challenge in computer vision, with approaches such as Neural Radiance Fields (NeRF) and Gaussian Splatting achieving remarkable progress. While Gaussian Splatting demonstrates strong performance…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Alexander Valverde , Brian Xu , Yuyin Zhou , Meng Xu , Hongyun Wang

Realtime 4D reconstruction for dynamic scenes remains a crucial challenge for autonomous driving perception. Most existing methods rely on depth estimation through self-supervision or multi-modality sensor fusion. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Xin Fei , Wenzhao Zheng , Yueqi Duan , Wei Zhan , Masayoshi Tomizuka , Kurt Keutzer , Jiwen Lu

Multi-view image diffusion models have significantly advanced open-domain 3D object generation. However, most existing models rely on 2D network architectures that lack inherent 3D biases, resulting in compromised geometric consistency. To…

计算机视觉与模式识别 · 计算机科学 2025-02-21 Hansheng Chen , Bokui Shen , Yulin Liu , Ruoxi Shi , Linqi Zhou , Connor Z. Lin , Jiayuan Gu , Hao Su , Gordon Wetzstein , Leonidas Guibas

Recent efforts in Gaussian-Splat-based Novel View Synthesis can achieve photorealistic rendering; however, such capability is limited in sparse-view scenarios due to sparse initialization and over-fitting floaters. Recent progress in depth…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Yutao Tang , Yuxiang Guo , Deming Li , Cheng Peng

Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or limited effective memory in recurrent designs. We present LoGeR…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Junyi Zhang , Charles Herrmann , Junhwa Hur , Chen Sun , Ming-Hsuan Yang , Forrester Cole , Trevor Darrell , Deqing Sun

Image Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Vincent Leroy , Yohann Cabon , Jérôme Revaud