English
Related papers

Related papers: GeoLRM: Geometry-Aware Large Reconstruction Model …

200 papers

We introduce GRM, a large-scale reconstructor capable of recovering a 3D asset from sparse-view images in around 0.1s. GRM is a feed-forward transformer-based model that efficiently incorporates multi-view information to translate the input…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Yinghao Xu , Zifan Shi , Wang Yifan , Hansheng Chen , Ceyuan Yang , Sida Peng , Yujun Shen , Gordon Wetzstein

Single-image 3D reconstruction with large reconstruction models (LRMs) has advanced rapidly, yet reconstructions often exhibit geometric inconsistencies and misaligned details that limit fidelity. We introduce GeoFusionLRM, a geometry-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Ahmet Burak Yildirim , Tuna Saygin , Duygu Ceylan , Aysegul Dundar

We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture;…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Kai Zhang , Sai Bi , Hao Tan , Yuanbo Xiangli , Nanxuan Zhao , Kalyan Sunkavalli , Zexiang Xu

We propose Long-LRM, a feed-forward 3D Gaussian reconstruction model for instant, high-resolution, 360{\deg} wide-coverage, scene-level reconstruction. Specifically, it takes in 32 input images at a resolution of 960x540 and produces the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chen Ziwen , Hao Tan , Kai Zhang , Sai Bi , Fujun Luan , Yicong Hong , Li Fuxin , Zexiang Xu

Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explicit 3D representations, such as 3D Gaussian splatting, has attracted significant attention…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Gyeongjin Kang , Seungtae Nam , Seungkwon Yang , Xiangyu Sun , Sameh Khamis , Abdelrahman Mohamed , Eunbyung Park

Despite recent advancements in the Large Reconstruction Model (LRM) demonstrating impressive results, when extending its input from single image to multiple images, it exhibits inefficiencies, subpar geometric and texture quality, as well…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Mengfei Li , Xiaoxiao Long , Yixun Liang , Weiyu Li , Yuan Liu , Peng Li , Wenhan Luo , Wenping Wang , Yike Guo

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ziqiao Ma , Xuweiyi Chen , Shoubin Yu , Sai Bi , Kai Zhang , Chen Ziwen , Sihan Xu , Jianing Yang , Zexiang Xu , Kalyan Sunkavalli , Mohit Bansal , Joyce Chai , Hao Tan

Feed-forward 3D generative models like the Large Reconstruction Model (LRM) have demonstrated exceptional generation speed. However, the transformer-based methods do not leverage the geometric priors of the triplane component in their…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Zhengyi Wang , Yikai Wang , Yifei Chen , Chendong Xiang , Shuo Chen , Dajiang Yu , Chongxuan Li , Hang Su , Jun Zhu

Recent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, photorealistic scene reconstruction. However, conventional 3DGS frameworks typically rely on sparse point clouds derived from Structure-from-Motion (SfM), which…

Graphics · Computer Science 2026-03-25 Yan Fang , Jianfei Ge , Jiangjian Xiao

Recent advances in generalizable Gaussian splatting (GS) have enabled feed-forward reconstruction of scenes from tens of input views. Long-LRM notably scales this paradigm to 32 input images at $950\times540$ resolution, achieving 360{\deg}…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Chen Ziwen , Hao Tan , Peng Wang , Zexiang Xu , Li Fuxin

We propose MeshLRM, a novel LRM-based approach that can reconstruct a high-quality mesh from merely four input images in less than one second. Different from previous large reconstruction models (LRMs) that focus on NeRF-based…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Xinyue Wei , Kai Zhang , Sai Bi , Hao Tan , Fujun Luan , Valentin Deschaintre , Kalyan Sunkavalli , Hao Su , Zexiang Xu

We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Zhengqin Li , Dilin Wang , Ka Chen , Zhaoyang Lv , Thu Nguyen-Phuoc , Milim Lee , Jia-Bin Huang , Lei Xiao , Cheng Zhang , Yufeng Zhu , Carl S. Marshall , Yufeng Ren , Richard Newcombe , Zhao Dong

We present recurrent geometry-aware neural networks that integrate visual information across multiple views of a scene into 3D latent feature tensors, while maintaining an one-to-one mapping between 3D physical locations in the world scene…

Computer Vision and Pattern Recognition · Computer Science 2018-11-15 Ricson Cheng , Ziyan Wang , Katerina Fragkiadaki

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Chen Wang , Hao Tan , Wang Yifan , Zhiqin Chen , Yuheng Liu , Kalyan Sunkavalli , Sai Bi , Lingjie Liu , Yiwei Hu

Generation of simulated detector response to collision products is crucial to data analysis in particle physics, but computationally very expensive. One subdetector, the calorimeter, dominates the computational time due to the high…

Instrumentation and Detectors · Physics 2023-11-16 Junze Liu , Aishik Ghosh , Dylan Smith , Pierre Baldi , Daniel Whiteson

We propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Tianyuan Zhang , Zhengfei Kuang , Haian Jin , Zexiang Xu , Sai Bi , Hao Tan , He Zhang , Yiwei Hu , Milos Hasan , William T. Freeman , Kai Zhang , Fujun Luan

3D Morphable Models (3DMMs) enable controllable facial geometry and expression editing for reconstruction, animation, and AR/VR, but traditional PCA-based mesh models are limited in resolution, detail, and photorealism. Neural volumetric…

Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Shihua Zhang , Qiuhong Shen , Shizun Wang , Tianbo Pan , Xinchao Wang

The increasing demand for 3D assets across various industries necessitates efficient and automated methods for 3D content creation. Leveraging 3D Gaussian Splatting, recent large reconstruction models (LRMs) have demonstrated the ability to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jingrui Ye , Lingting Zhu , Runze Zhang , Zeyu Hu , Yingda Yin , Lanjiong Li , Lequan Yu , Qingmin Liao

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen
‹ Prev 1 2 3 10 Next ›