English
Related papers

Related papers: GRM: Large Gaussian Reconstruction Model for Effic…

200 papers

We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture;…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Kai Zhang , Sai Bi , Hao Tan , Yuanbo Xiangli , Nanxuan Zhao , Kalyan Sunkavalli , Zexiang Xu

Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explicit 3D representations, such as 3D Gaussian splatting, has attracted significant attention…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Gyeongjin Kang , Seungtae Nam , Seungkwon Yang , Xiangyu Sun , Sameh Khamis , Abdelrahman Mohamed , Eunbyung Park

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only 11 GB GPU memory. Previous works neglect the inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Chubin Zhang , Hongliang Song , Yi Wei , Yu Chen , Jiwen Lu , Yansong Tang

We propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Tianyuan Zhang , Zhengfei Kuang , Haian Jin , Zexiang Xu , Sai Bi , Hao Tan , He Zhang , Yiwei Hu , Milos Hasan , William T. Freeman , Kai Zhang , Fujun Luan

We propose Long-LRM, a feed-forward 3D Gaussian reconstruction model for instant, high-resolution, 360{\deg} wide-coverage, scene-level reconstruction. Specifically, it takes in 32 input images at a resolution of 960x540 and produces the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chen Ziwen , Hao Tan , Kai Zhang , Sai Bi , Fujun Luan , Yicong Hong , Li Fuxin , Zexiang Xu

We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Zhengqin Li , Dilin Wang , Ka Chen , Zhaoyang Lv , Thu Nguyen-Phuoc , Milim Lee , Jia-Bin Huang , Lei Xiao , Cheng Zhang , Yufeng Zhu , Carl S. Marshall , Yufeng Ren , Richard Newcombe , Zhao Dong

The increasing demand for 3D assets across various industries necessitates efficient and automated methods for 3D content creation. Leveraging 3D Gaussian Splatting, recent large reconstruction models (LRMs) have demonstrated the ability to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jingrui Ye , Lingting Zhu , Runze Zhang , Zeyu Hu , Yingda Yin , Lanjiong Li , Lequan Yu , Qingmin Liao

Computed Tomography serves as an indispensable tool in clinical workflows, providing non-invasive visualization of internal anatomical structures. Existing CT reconstruction works are limited to small-capacity model architecture and…

Image and Video Processing · Electrical Eng. & Systems 2025-05-27 Yifan Liu , Wuyang Li , Weihao Yu , Chenxin Li , Alexandre Alahi , Max Meng , Yixuan Yuan

Feed-forward 3D generative models like the Large Reconstruction Model (LRM) have demonstrated exceptional generation speed. However, the transformer-based methods do not leverage the geometric priors of the triplane component in their…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Zhengyi Wang , Yikai Wang , Yifei Chen , Chendong Xiang , Shuo Chen , Dajiang Yu , Chongxuan Li , Hang Su , Jun Zhu

3D content creation has achieved significant progress in terms of both quality and speed. Although current feed-forward models can produce 3D objects in seconds, their resolution is constrained by the intensive computation required during…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Jiaxiang Tang , Zhaoxi Chen , Xiaokang Chen , Tengfei Wang , Gang Zeng , Ziwei Liu

Modeling 3D articulated objects with realistic geometry, textures, and kinematics is essential for a wide range of applications. However, existing optimization-based reconstruction methods often require dense multi-view inputs and expensive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Sylvia Yuan , Ruoxi Shi , Xinyue Wei , Xiaoshuai Zhang , Hao Su , Minghua Liu

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction…

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Jiawei Ren , Kevin Xie , Ashkan Mirzaei , Hanxue Liang , Xiaohui Zeng , Karsten Kreis , Ziwei Liu , Antonio Torralba , Sanja Fidler , Seung Wook Kim , Huan Ling

We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Yicong Hong , Kai Zhang , Jiuxiang Gu , Sai Bi , Yang Zhou , Difan Liu , Feng Liu , Kalyan Sunkavalli , Trung Bui , Hao Tan

Despite recent advancements in the Large Reconstruction Model (LRM) demonstrating impressive results, when extending its input from single image to multiple images, it exhibits inefficiencies, subpar geometric and texture quality, as well…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Mengfei Li , Xiaoxiao Long , Yixun Liang , Weiyu Li , Yuan Liu , Peng Li , Wenhan Luo , Wenping Wang , Yike Guo

We present Real-time Gaussian SLAM (RTG-SLAM), a real-time 3D reconstruction system with an RGBD camera for large-scale environments using Gaussian splatting. The system features a compact Gaussian representation and a highly efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Zhexi Peng , Tianjia Shao , Yong Liu , Jingke Zhou , Yin Yang , Jingdong Wang , Kun Zhou

Accurate 3D reconstruction of vehicles is vital for applications such as vehicle inspection, predictive maintenance, and urban planning. Existing methods like Neural Radiance Fields and Gaussian Splatting have shown impressive results but…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Davide Di Nucci , Matteo Tomei , Guido Borghi , Luca Ciuffreda , Roberto Vezzani , Rita Cucchiara

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ziqiao Ma , Xuweiyi Chen , Shoubin Yu , Sai Bi , Kai Zhang , Chen Ziwen , Sihan Xu , Jianing Yang , Zexiang Xu , Kalyan Sunkavalli , Mohit Bansal , Joyce Chai , Hao Tan

We aim to address sparse-view reconstruction of a 3D scene by leveraging priors from large-scale vision models. While recent advancements such as 3D Gaussian Splatting (3DGS) have demonstrated remarkable successes in 3D reconstruction,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Hanyang Yu , Xiaoxiao Long , Ping Tan

3D Morphable Models (3DMMs) enable controllable facial geometry and expression editing for reconstruction, animation, and AR/VR, but traditional PCA-based mesh models are limited in resolution, detail, and photorealism. Neural volumetric…

‹ Prev 1 2 3 10 Next ›