English
Related papers

Related papers: Long-LRM: Long-sequence Large Reconstruction Model…

200 papers

We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture;…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Kai Zhang , Sai Bi , Hao Tan , Yuanbo Xiangli , Nanxuan Zhao , Kalyan Sunkavalli , Zexiang Xu

Recent advances in generalizable Gaussian splatting (GS) have enabled feed-forward reconstruction of scenes from tens of input views. Long-LRM notably scales this paradigm to 32 input images at $950\times540$ resolution, achieving 360{\deg}…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Chen Ziwen , Hao Tan , Peng Wang , Zexiang Xu , Li Fuxin

We introduce GRM, a large-scale reconstructor capable of recovering a 3D asset from sparse-view images in around 0.1s. GRM is a feed-forward transformer-based model that efficiently incorporates multi-view information to translate the input…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Yinghao Xu , Zifan Shi , Wang Yifan , Hansheng Chen , Ceyuan Yang , Sida Peng , Yujun Shen , Gordon Wetzstein

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only 11 GB GPU memory. Previous works neglect the inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Chubin Zhang , Hongliang Song , Yi Wei , Yu Chen , Jiwen Lu , Yansong Tang

Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explicit 3D representations, such as 3D Gaussian splatting, has attracted significant attention…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Gyeongjin Kang , Seungtae Nam , Seungkwon Yang , Xiangyu Sun , Sameh Khamis , Abdelrahman Mohamed , Eunbyung Park

Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ziqiao Ma , Xuweiyi Chen , Shoubin Yu , Sai Bi , Kai Zhang , Chen Ziwen , Sihan Xu , Jianing Yang , Zexiang Xu , Kalyan Sunkavalli , Mohit Bansal , Joyce Chai , Hao Tan

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction…

3D content creation has achieved significant progress in terms of both quality and speed. Although current feed-forward models can produce 3D objects in seconds, their resolution is constrained by the intensive computation required during…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Jiaxiang Tang , Zhaoxi Chen , Xiaokang Chen , Tengfei Wang , Gang Zeng , Ziwei Liu

We propose MeshLRM, a novel LRM-based approach that can reconstruct a high-quality mesh from merely four input images in less than one second. Different from previous large reconstruction models (LRMs) that focus on NeRF-based…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Xinyue Wei , Kai Zhang , Sai Bi , Hao Tan , Fujun Luan , Valentin Deschaintre , Kalyan Sunkavalli , Hao Su , Zexiang Xu

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Chen Wang , Hao Tan , Wang Yifan , Zhiqin Chen , Yuheng Liu , Kalyan Sunkavalli , Sai Bi , Lingjie Liu , Yiwei Hu

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Jiawei Ren , Kevin Xie , Ashkan Mirzaei , Hanxue Liang , Xiaohui Zeng , Karsten Kreis , Ziwei Liu , Antonio Torralba , Sanja Fidler , Seung Wook Kim , Huan Ling

Sparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity of large-scale training datasets and the absence of direct…

Image and Video Processing · Electrical Eng. & Systems 2026-01-29 Guofeng Zhang , Ruyi Zha , Hao He , Yixun Liang , Alan Yuille , Hongdong Li , Yuanhao Cai

We propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Tianyuan Zhang , Zhengfei Kuang , Haian Jin , Zexiang Xu , Sai Bi , Hao Tan , He Zhang , Yiwei Hu , Milos Hasan , William T. Freeman , Kai Zhang , Fujun Luan

We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Zhengqin Li , Dilin Wang , Ka Chen , Zhaoyang Lv , Thu Nguyen-Phuoc , Milim Lee , Jia-Bin Huang , Lei Xiao , Cheng Zhang , Yufeng Zhu , Carl S. Marshall , Yufeng Ren , Richard Newcombe , Zhao Dong

Complete reconstruction of surgical scenes is crucial for robot-assisted surgery (RAS). Deep depth estimation is promising but existing works struggle with depth discontinuities, resulting in noisy predictions at object boundaries and do…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Xu Wang , Shuai Zhang , Baoru Huang , Danail Stoyanov , Evangelos B. Mazomenos

Despite recent advancements in the Large Reconstruction Model (LRM) demonstrating impressive results, when extending its input from single image to multiple images, it exhibits inefficiencies, subpar geometric and texture quality, as well…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Mengfei Li , Xiaoxiao Long , Yixun Liang , Weiyu Li , Yuan Liu , Peng Li , Wenhan Luo , Wenping Wang , Yike Guo

We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Yicong Hong , Kai Zhang , Jiuxiang Gu , Sai Bi , Yang Zhou , Difan Liu , Feng Liu , Kalyan Sunkavalli , Trung Bui , Hao Tan

We present 3DGS-LM, a new method that accelerates the reconstruction of 3D Gaussian Splatting (3DGS) by replacing its ADAM optimizer with a tailored Levenberg-Marquardt (LM). Existing methods reduce the optimization time by decreasing the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Lukas Höllein , Aljaž Božič , Michael Zollhöfer , Matthias Nießner

3D Gaussian Splatting (3DGS) is an increasingly popular novel view synthesis approach due to its fast rendering time, and high-quality output. However, scaling 3DGS to large (or intricate) scenes is challenging due to its large memory…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Hexu Zhao , Xiwen Min , Xiaoteng Liu , Moonjun Gong , Yiming Li , Ang Li , Saining Xie , Jinyang Li , Aurojit Panda

Radiance field methods have achieved photorealistic novel view synthesis and geometry reconstruction. But they are mostly applied in per-scene optimization or small-baseline settings. While several recent works investigate feed-forward…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Anpei Chen , Haofei Xu , Stefano Esposito , Siyu Tang , Andreas Geiger
‹ Prev 1 2 3 10 Next ›