中文
相关论文

相关论文: Mem3R: Streaming 3D Reconstruction with Hybrid Mem…

200 篇论文

The recent paradigm shift in 3D vision led to the rise of foundation models with remarkable capabilities in 3D perception from uncalibrated images. However, extending these models to large-scale RGB stream 3D reconstruction remains…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Leonid Antsfeld , Boris Chidlovskii , Yohann Cabon , Vincent Leroy , Jerome Revaud

Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Chong Cheng , Peilin Tao , Nanjie Yao , Guanzhi Ding , Xianda Chen , Yuansen Du , Xiaoyang Guo , Wei Yin , Weiqiang Ren , Qian Zhang , Zhengqing Chen , Hao Wang

We present Human3R, a unified, feed-forward framework for online 4D human-scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Yue Chen , Xingyu Chen , Yuxuan Xue , Anpei Chen , Yuliang Xiu , Gerard Pons-Moll

Despite recent advances in feed-forward 3D Gaussian Splatting, generalizable 3D reconstruction remains challenging, particularly in multi-view correspondence modeling. Existing approaches face a fundamental trade-off: explicit methods…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Heng Jia , Linchao Zhu , Na Zhao

Streaming visual transformers like StreamVGGT achieve strong 3D perception but suffer from unbounded growth of key value (KV) memory, which limits scalability. We propose a training-free, inference-time token eviction policy that bounds…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Soroush Mahdi , Fardin Ayar , Ehsan Javanmardi , Manabu Tsukada , Mahdi Javanmardi

Streaming feed-forward 3D reconstruction enables real-time joint estimation of scene geometry and camera poses from RGB images. However, without explicit dynamic reasoning, streaming models can be affected by moving objects, causing…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Feiran Wang , Zezhou Shang , Gaowen Liu , Yan Yan

Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Dong Zhuo , Wenzhao Zheng , Jiahe Guo , Yuqi Wu , Jie Zhou , Jiwen Lu

The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neural implicit 3D representation methods face significant…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Luoxi Zhang , Pragyan Shrestha , Yu Zhou , Chun Xie , Itaru Kitahara

Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Congrong Xu , Huachen Gao , Xingyu Chen , Yuliang Xiu , Jun Gao , Anpei Chen

We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Zizun Li , Jianjun Zhou , Yifan Wang , Haoyu Guo , Wenzheng Chang , Yang Zhou , Haoyi Zhu , Junyi Chen , Chunhua Shen , Tong He

Transformer-based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi-view observations, offering strong performance across challenging visual conditions. As these models scale to larger…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Haoyu Zhang , Zeyu Zhang , Zedong Zhou , Yang Zhao , Hao Tang

Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across five benchmarks and discover a structural bottleneck: the…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Kejun Ren , Lei Jin , Tianxin Huang , Lianming Xu , Li Wang

Dense matching methods like DUSt3R regress pairwise pointmaps for 3D reconstruction. However, the reliance on pairwise prediction and the limited generalization capability inherently restrict the global geometric consistency. In this work,…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Yuheng Yuan , Qiuhong Shen , Shizun Wang , Xingyi Yang , Xinchao Wang

Image matching is a key component of modern 3D vision algorithms, essential for accurate scene reconstruction and localization. MASt3R redefines image matching as a 3D task by leveraging DUSt3R and introducing a fast reciprocal matching…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Jingxing Li , Yongjae Lee , Abhay Kumar Yadav , Cheng Peng , Rama Chellappa , Deliang Fan

Processing long videos with multimodal large language models (MLLMs) poses a significant computational challenge, as the model's self-attention mechanism scales quadratically with the number of video tokens, resulting in high computational…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Kaibin Wang , Mingbao Lin

We present Online3R, a new sequential reconstruction framework that is capable of adapting to new scenes through online learning, effectively resolving inconsistency issues. Specifically, we introduce a set of learnable lightweight visual…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Shunkai Zhou , Zike Yan , Fei Xue , Dong Wu , Yuchen Deng , Hongbin Zha

In this paper, we introduce SLAM3R, a novel and effective system for real-time, high-quality, dense 3D reconstruction using RGB videos. SLAM3R provides an end-to-end solution by seamlessly integrating local 3D reconstruction and global…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Yuzheng Liu , Siyan Dong , Shuzhe Wang , Yingda Yin , Yanchao Yang , Qingnan Fan , Baoquan Chen

Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video inputs with hundreds or thousands of frames remains…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Zecheng Tang , Jiaye Fu , Qiankun Gao , Haijie Li , Yanmin Wu , Jiaqi Zhang , Siwei Ma , Jian Zhang

Fast 3D clothed human reconstruction from monocular video remains a significant challenge in computer vision, particularly in balancing computational efficiency with reconstruction quality. Current approaches are either focused on static…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Matthew Marchellus , Nadhira Noor , In Kyu Park

We present AMB3R, a multi-view feed-forward model for dense 3D reconstruction on a metric-scale that addresses diverse 3D vision tasks. The key idea is to leverage a sparse, yet compact, volumetric scene representation as our backend,…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Hengyi Wang , Lourdes Agapito