English
Related papers

Related papers: Mem3R: Streaming 3D Reconstruction with Hybrid Mem…

200 papers

Building a robust perception module is crucial for visuomotor policy learning. While recent methods incorporate pre-trained 2D foundation models into robotic perception modules to leverage their strong semantic understanding, they struggle…

Robotics · Computer Science 2025-07-14 Wenbo Cui , Chengyang Zhao , Yuhui Chen , Haoran Li , Zhizheng Zhang , Dongbin Zhao , He Wang

Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for image recognition problems. Nevertheless, it is not trivial when utilizing a CNN for learning spatio-temporal video representation. A few studies have…

Computer Vision and Pattern Recognition · Computer Science 2017-11-29 Zhaofan Qiu , Ting Yao , Tao Mei

Imitation Learning can train robots to perform complex and diverse manipulation tasks, but learned policies are brittle with observations outside of the training distribution. 3D scene representations that incorporate observations from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Albert Wilcox , Mohamed Ghanem , Masoud Moghani , Pierre Barroso , Benjamin Joffe , Animesh Garg

Continual learning enables models to acquire new knowledge over time while retaining previously learned capabilities. However, its application to text-to-3D generation remains unexplored. We present ReConText3D, the first framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Muhammad Ahmed Ullah Khan , Muhammad Haris Bin Amir , Didier Stricker , Muhammad Zeshan Afzal

Simultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems. To achieve this, recent approaches resort to 2D-to-3D feature alignment paradigm, which leads to limited 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Qi Xu , Dongxu Wei , Lingzhe Zhao , Wenpu Li , Zhangchi Huang , Shunping Ji , Peidong Liu

For Embodied AI, jointly reconstructing dynamic hands and the dense scene context is crucial for understanding physical interaction. However, most existing methods recover isolated hands in local coordinates, overlooking the surrounding 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Wendi Hu , Haonan Zhou , Wenhao Hu , Gaoang Wang

Recent feed-forward reconstruction models like VGGT and $\pi^3$ achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, limiting their practical deployment. While existing streaming…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Tianye Ding , Yiming Xie , Yiqing Liang , Moitreya Chatterjee , Pedro Miraldo , Huaizu Jiang

3D reconstruction, which aims to recover the dense three-dimensional structure of a scene, is a cornerstone technology for numerous applications, including augmented/virtual reality, autonomous driving, and robotics. While traditional…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Wei Zhang , Yihang Wu , Songhua Li , Wenjie Ma , Xin Ma , Qiang Li , Qi Wang

Robust object tracking requires knowledge and understanding of the object being tracked: its appearance, its motion, and how it changes over time. A tracker must be able to modify its underlying model and adapt to new observations. We…

Computer Vision and Pattern Recognition · Computer Science 2018-02-28 Daniel Gordon , Ali Farhadi , Dieter Fox

Current multi-view 3D reconstruction methods rely on accurate camera calibration and pose estimation, requiring complex and time-intensive pre-processing that hinders their practical deployment. To address this challenge, we introduce…

Graphics · Computer Science 2025-08-07 Haodong Zhu , Changbai Li , Yangyang Ren , Zichao Feng , Xuhui Liu , Hanlin Chen , Xiantong Zhen , Baochang Zhang

We present a novel method to learn temporally consistent 3D reconstruction of clothed people from a monocular video. Recent methods for 3D human reconstruction from monocular video using volumetric, implicit or parametric human shape…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Akin Caliskan , Armin Mustafa , Adrian Hilton

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Sara Rojas , Matthieu Armando , Bernard Ghamen , Philippe Weinzaepfel , Vincent Leroy , Gregory Rogez

Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, e.g., convolutional neural networks (CNNs) and…

Computer Vision and Pattern Recognition · Computer Science 2016-11-18 Junbo Wang , Wei Wang , Yan Huang , Liang Wang , Tieniu Tan

Multi-view 3D reconstruction has remained an essential yet challenging problem in the field of computer vision. While DUSt3R and its successors have achieved breakthroughs in 3D reconstruction from unposed images, these methods exhibit…

Image and Video Processing · Electrical Eng. & Systems 2025-09-16 Sidun Liu , Wenyu Li , Peng Qiao , Yong Dou

Long-term temporal fusion is a crucial but often overlooked technique in camera-based Bird's-Eye-View (BEV) 3D perception. Existing methods are mostly in a parallel manner. While parallel fusion can benefit from long-term information, it…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Chunrui Han , Jinrong Yang , Jianjian Sun , Zheng Ge , Runpei Dong , Hongyu Zhou , Weixin Mao , Yuang Peng , Xiangyu Zhang

Online 3D reconstruction from streaming inputs requires both long-term temporal consistency and efficient memory usage. Although causal variants of VGGT address this challenge through a key-value (KV) cache mechanism, the cache grows…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Runze Wang , Yuxuan Song , Youcheng Cai , Ligang Liu

Accurate and computationally efficient 3D medical image segmentation remains a critical challenge in clinical workflows. Transformer-based architectures often demonstrate superior global contextual modeling but at the expense of excessive…

Image and Video Processing · Electrical Eng. & Systems 2026-02-19 Kavyansh Tyagi , Vishwas Rathi , Puneet Goyal

We propose a transformer-based neural network architecture for multi-object 3D reconstruction from RGB videos. It relies on two alternative ways to represent its knowledge: as a global 3D grid of features and an array of view-specific 2D…

Computer Vision and Pattern Recognition · Computer Science 2022-08-29 Michał J. Tyszkiewicz , Kevis-Kokitsi Maninis , Stefan Popov , Vittorio Ferrari

In recent years, 3D visual foundation models pioneered by pointmap-based approaches such as DUSt3R have attracted a lot of interest, achieving impressive accuracy and strong generalization across diverse scenes. However, these methods are…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shuang Guo , Filbert Febryanto , Lei Sun , Guillermo Gallego

DUSt3R-based end-to-end scene reconstruction has recently shown promising results in dense visual SLAM. However, most existing methods only use image pairs to estimate pointmaps, overlooking spatial memory and global consistency.To this…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Guole Shen , Tianchen Deng , Yanbo Wang , Yongtao Chen , Yilin Shen , Jiuming Liu , Jingchuan Wang