中文
相关论文

相关论文: Long-LRM++: Preserving Fine Details in Feed-Forwar…

200 篇论文

We propose Long-LRM, a feed-forward 3D Gaussian reconstruction model for instant, high-resolution, 360{\deg} wide-coverage, scene-level reconstruction. Specifically, it takes in 32 input images at a resolution of 960x540 and produces the…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Chen Ziwen , Hao Tan , Kai Zhang , Sai Bi , Fujun Luan , Yicong Hong , Li Fuxin , Zexiang Xu

Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explicit 3D representations, such as 3D Gaussian splatting, has attracted significant attention…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Gyeongjin Kang , Seungtae Nam , Seungkwon Yang , Xiangyu Sun , Sameh Khamis , Abdelrahman Mohamed , Eunbyung Park

We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture;…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Kai Zhang , Sai Bi , Hao Tan , Yuanbo Xiangli , Nanxuan Zhao , Kalyan Sunkavalli , Zexiang Xu

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction…

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability.…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Chen Wang , Hao Tan , Wang Yifan , Zhiqin Chen , Yuheng Liu , Kalyan Sunkavalli , Sai Bi , Lingjie Liu , Yiwei Hu

Capturing relightable 3D assets from real-world objects is a widely researched problem. Several per-scene optimization-based methods, based on 3D Gaussian splatting (3DGS), support relighting; however, they usually require dense input…

图形学 · 计算机科学 2026-05-29 Guangming Fu , Jiahui Fan , Jian Yang , Miloš Hašan , Beibei Wang

We aim to address sparse-view reconstruction of a 3D scene by leveraging priors from large-scale vision models. While recent advancements such as 3D Gaussian Splatting (3DGS) have demonstrated remarkable successes in 3D reconstruction,…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Hanyang Yu , Xiaoxiao Long , Ping Tan

We propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Tianyuan Zhang , Zhengfei Kuang , Haian Jin , Zexiang Xu , Sai Bi , Hao Tan , He Zhang , Yiwei Hu , Milos Hasan , William T. Freeman , Kai Zhang , Fujun Luan

Implicit neural representations (INRs) have achieved remarkable success in image representation and compression, but they require substantial training time and memory. Meanwhile, recent 2D Gaussian Splatting (GS) methods (\textit{e.g.},…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Tiantian Li , Xinjie Zhang , Xingtong Ge , Tongda Xu , Dailan He , Jun Zhang , Yan Wang

Complete reconstruction of surgical scenes is crucial for robot-assisted surgery (RAS). Deep depth estimation is promising but existing works struggle with depth discontinuities, resulting in noisy predictions at object boundaries and do…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Xu Wang , Shuai Zhang , Baoru Huang , Danail Stoyanov , Evangelos B. Mazomenos

Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details.…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Xinyuan Hu , Changyue Shi , Chuxiao Yang , Minghao Chen , Jiajun Ding , Tao Wei , Chen Wei , Zhou Yu , Min Tan

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only 11 GB GPU memory. Previous works neglect the inherent…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Chubin Zhang , Hongliang Song , Yi Wei , Yu Chen , Jiwen Lu , Yansong Tang

As the Large Language Model (LLM) becomes increasingly important in various domains. However, the following challenges still remain unsolved in accelerating LLM inference: (1) Synchronized partial softmax update. The softmax operation…

机器学习 · 计算机科学 2024-01-08 Ke Hong , Guohao Dai , Jiaming Xu , Qiuli Mao , Xiuhong Li , Jun Liu , Kangdi Chen , Yuhan Dong , Yu Wang

We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows impacts feed-forward 3D reconstruction. Although recent object-centric feed-forward methods deliver robust, high-quality reconstruction,…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Zhengqin Li , Cheng Zhang , Jakob Engel , Zhao Dong

Incrementally recovering real-sized 3D geometry from a pose-free RGB stream is a challenging task in 3D reconstruction, requiring minimal assumptions on input data. Existing methods can be broadly categorized into end-to-end and visual…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Linqing Zhao , Xiuwei Xu , Yirui Wang , Hao Wang , Wenzhao Zheng , Yansong Tang , Haibin Yan , Jiwen Lu

The advent of 3D Gaussian Splatting (3D-GS) techniques and their dynamic scene modeling variants, 4D-GS, offers promising prospects for real-time rendering of dynamic surgical scenarios. However, the prerequisite for modeling dynamic scenes…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Hengyu Liu , Yifan Liu , Chenxin Li , Wuyang Li , Yixuan Yuan

Existing feed-forward 3D Gaussian Splatting methods predict pixel-aligned primitives, leading to a quadratic growth in primitive count as resolution increases. This fundamentally limits their scalability, making high-resolution synthesis…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yixing Lao , Xuyang Bai , Xiaoyang Wu , Nuoyuan Yan , Zixin Luo , Tian Fang , Jean-Daniel Nahmias , Yanghai Tsin , Shiwei Li , Hengshuang Zhao

3D super-resolution (3DSR) aims to reconstruct high-resolution (HR) 3D scenes from low-resolution (LR) multi-view images. Existing methods rely on dense LR inputs and per-scene optimization, which restricts the high-frequency priors for…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Feng , Xiangbo Wang , Tieshi Zhong , Chengkai Wang , Yiting Zhao , Tianxiang Xu , Zhenzhong Kuang , Feiwei Qin , Xuefei Yin , Yanming Zhu

Real-time 3D reconstruction is crucial for robotics and augmented reality, yet current simultaneous localization and mapping(SLAM) approaches often struggle to maintain structural consistency and robust pose estimation in the presence of…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Xu Wang , Boyao Han , Xiaojun Chen , Ying Liu , Ruihui Li

The recent 3D Gaussian splatting (3D-GS) has shown remarkable rendering fidelity and efficiency compared to NeRF-based neural scene representations. While demonstrating the potential for real-time rendering, 3D-GS encounters rendering…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Kerui Ren , Lihan Jiang , Tao Lu , Mulin Yu , Linning Xu , Zhangkai Ni , Bo Dai
‹ 上一页 1 2 3 10 下一页 ›