English
Related papers

Related papers: Mem3R: Streaming 3D Reconstruction with Hybrid Mem…

200 papers

In this paper, we propose a new framework for online 3D scene perception. Conventional 3D scene perception methods are offline, i.e., take an already reconstructed 3D scene geometry as input, which is not applicable in robotic applications…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Xiuwei Xu , Chong Xia , Ziwei Wang , Linqing Zhao , Yueqi Duan , Jie Zhou , Jiwen Lu

In this paper, we propose a long-sequence modeling framework, named StreamPETR, for multi-view 3D object detection. Built upon the sparse query design in the PETR series, we systematically develop an object-centric temporal mechanism. The…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Shihao Wang , Yingfei Liu , Tiancai Wang , Ying Li , Xiangyu Zhang

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Chen Wang , Hao Tan , Wang Yifan , Zhiqin Chen , Yuheng Liu , Kalyan Sunkavalli , Sai Bi , Lingjie Liu , Yiwei Hu

We present Wid3R, a feed-forward neural network for multi-view visual geometry reconstruction that supports wide field-of-view camera models. Unlike existing methods that assume rectified or pinhole inputs, Wid3R directly models wide-angle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Dongki Jung , Jaehoon Choi , Adil Qureshi , Somi Jeong , Dinesh Manocha , Suyong Yeon

Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets. In contrast, the limited…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Xingyu Chen , Yue Chen , Yuliang Xiu , Andreas Geiger , Anpei Chen

Recent trends in sparse-view 3D reconstruction have taken two different paths: feed-forward reconstruction that predicts pixel-aligned point maps without a complete geometry, and generative 3D reconstruction that generates complete geometry…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Siyou Lin , Zhou Xue , Hongwen Zhang , Liang An , Dongping Li , Shaohui Jiao , Yebin Liu

Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attention leads to quadratic compute complexity in sequence length and memory usage due to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Kunyang Li , Mubarak Shah , Yuzhang Shang

We introduce a novel, data-driven approach for reconstructing temporally coherent 3D motion from unstructured and potentially partial observations of non-rigidly deforming shapes. Our goal is to achieve high-fidelity motion reconstructions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Aymen Merrouche , Stefanie Wuhrer , Edmond Boyer

Traditionally, creating photo-realistic 3D head avatars requires a studio-level multi-view capture setup and expensive optimization during test-time, limiting the use of digital human doubles to the VFX industry or offline renderings. To…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Tobias Kirschstein , Javier Romero , Artem Sevastopolsky , Matthias Nießner , Shunsuke Saito

This paper presents a vector HD-mapping algorithm that formulates the mapping as a tracking task and uses a history of memory latents to ensure consistent reconstructions over time. Our method, MapTracker, accumulates a sensor stream into…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Jiacheng Chen , Yuefan Wu , Jiaqi Tan , Hang Ma , Yasutaka Furukawa

Many learning tasks involve multi-modal data streams, where continuous data from different modes convey a comprehensive description about objects. A major challenge in this context is how to efficiently interpret multi-modal information in…

Machine Learning · Computer Science 2020-07-24 Amila Silva , Shanika Karunasekera , Christopher Leckie , Ling Luo

Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging and benefits from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Jisu Nam , Jahyeok Koo , Soowon Son , Jaewoo Jung , Honggyu An , Junhwa Hur , Seungryong Kim

Despite the impressive progress of telepresence systems for room-scale scenes with static and dynamic scene entities, expanding their capabilities to scenarios with larger dynamic environments beyond a fixed size of a few square-meters…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Leif Van Holland , Patrick Stotko , Stefan Krumpen , Reinhard Klein , Michael Weinmann

The core challenge for streaming video generation is maintaining the content consistency in long context, which poses high requirement for the memory design. Most existing solutions maintain the memory by compressing historical frames with…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Sihui Ji , Xi Chen , Shuai Yang , Xin Tao , Pengfei Wan , Hengshuang Zhao

Hand-object 3D reconstruction has become increasingly important for applications in human-robot interaction and immersive AR/VR experiences. A common approach for object-agnostic hand-object reconstruction from RGB sequences involves a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Anilkumar Swamy , Vincent Leroy , Philippe Weinzaepfel , Jean-Sébastien Franco , Grégory Rogez

Multimodal large language models (MLLMs) have shown strong performance on offline video understanding, but most are limited to offline inference or have weak online reasoning, making multi-turn interaction over continuously arriving video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Lu Wang , Zhuoran Jin , Yupu Hao , Yubo Chen , Kang Liu , Yulong Ao , Jun Zhao

With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal…

Linear sequence modeling methods, such as linear attention, state space modeling, and linear RNNs, offer significant efficiency improvements by reducing the complexity of training and inference. However, these methods typically compress the…

Computation and Language · Computer Science 2025-11-19 Jusen Du , Weigao Sun , Disen Lan , Jiaxi Hu , Yu Cheng

Real-time 3D scene reconstruction from RGB-D sensor data, as well as the exploration of such data in VR/AR settings, has seen tremendous progress in recent years. The combination of both these components into telepresence systems, however,…

Human-Computer Interaction · Computer Science 2019-04-15 Patrick Stotko , Stefan Krumpen , Matthias B. Hullin , Michael Weinmann , Reinhard Klein

Mixed reality (MR) is a key technology which promises to change the future of warfare. An MR hybrid of physical outdoor environments and virtual military training will enable engagements with long distance enemies, both real and simulated.…

Image and Video Processing · Electrical Eng. & Systems 2022-11-23 David Ramirez , Suren Jayasuriya , Andreas Spanias
‹ Prev 1 3 4 5 6 7 10 Next ›