English
Related papers

Related papers: Scal3R: Scalable Test-Time Training for Large-Scal…

200 papers

Reconstructing detailed 3D scenes from single-view images remains a challenging task due to limitations in existing approaches, which primarily focus on geometric shape recovery, overlooking object appearances and fine shape details. To…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Yixin Chen , Junfeng Ni , Nan Jiang , Yaowei Zhang , Yixin Zhu , Siyuan Huang

Scene parsing is a technique that consist on giving a label to all pixels in an image according to the class they belong to. To ensure a good visual coherence and a high class accuracy, it is essential for a scene parser to capture image…

Computer Vision and Pattern Recognition · Computer Science 2013-06-13 Pedro H. O. Pinheiro , Ronan Collobert

Large kernel convolutions offer a scalable alternative to vision transformers for high-resolution 3D volumetric analysis, yet naively increasing kernel size often leads to optimization instability. Motivated by the spatial bias inherent in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Ho Hin Lee , Quan Liu , Shunxing Bao , Yuankai Huo , Bennett A. Landman

Despite the potential of neural scene representations to effectively compress 3D scalar fields at high reconstruction quality, the computational complexity of the training and data reconstruction step using scene representation networks…

Graphics · Computer Science 2022-07-26 Sebastian Weiss , Philipp Hermüller , Rüdiger Westermann

In this paper, we present a novel, scalable approach for constructing open set, instance-level 3D scene representations, advancing open world understanding of 3D environments. Existing methods require pre-constructed 3D scenes and face…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Rafay Mohiuddin , Sai Manoj Prakhya , Fiona Collins , Ziyuan Liu , André Borrmann

We introduce G-CUT3R, a novel feed-forward approach for guided 3D scene reconstruction that enhances the CUT3R model by integrating prior information. Unlike existing feed-forward methods that rely solely on input images, our method…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ramil Khafizov , Artem Komarichev , Ruslan Rakhimov , Peter Wonka , Evgeny Burnaev

Stylizing 3D scenes instantly while maintaining multi-view consistency and faithfully resembling a style image remains a significant challenge. Current state-of-the-art 3D stylization methods typically involve computationally intensive…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Peng Wang , Xiang Liu , Peidong Liu

Building a robust perception module is crucial for visuomotor policy learning. While recent methods incorporate pre-trained 2D foundation models into robotic perception modules to leverage their strong semantic understanding, they struggle…

Robotics · Computer Science 2025-07-14 Wenbo Cui , Chengyang Zhao , Yuhui Chen , Haoran Li , Zhizheng Zhang , Dongbin Zhao , He Wang

Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and interacting with the physical world. While traditional methods achieve high fidelity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Weijie Wang , Qihang Cao , Sensen Gao , Donny Y. Chen , Haofei Xu , Wenjing Bian , Songyou Peng , Tat-Jen Cham , Chuanxia Zheng , Andreas Geiger , Jianfei Cai , Jia-Wang Bian , Bohan Zhuang

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jiaxin Huang , Yuanbo Yang , Bangbang Yang , Lin Ma , Yuewen Ma , Yiyi Liao

We extend neural radiance fields (NeRFs) to dynamic large-scale urban scenes. Prior work tends to reconstruct single video clips of short durations (up to 10 seconds). Two reasons are that such methods (a) tend to scale linearly with the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Haithem Turki , Jason Y. Zhang , Francesco Ferroni , Deva Ramanan

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a prohibitive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Weining Ren , Xiao Tan , Kai Han

Recent advances in vision foundation models have revolutionized geometry reconstruction and semantic understanding. Yet, most of the existing approaches treat these capabilities in isolation, leading to redundant pipelines and compounded…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chaoyi Zhou , Run Wang , Feng Luo , Mert D. Pesé , Zhiwen Fan , Yiqi Zhong , Siyu Huang

In this paper, we propose a new framework for online 3D scene perception. Conventional 3D scene perception methods are offline, i.e., take an already reconstructed 3D scene geometry as input, which is not applicable in robotic applications…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Xiuwei Xu , Chong Xia , Ziwei Wang , Linqing Zhao , Yueqi Duan , Jie Zhou , Jiwen Lu

Accurate and robust 3D scene reconstruction from casual, in-the-wild videos can significantly simplify robot deployment to new environments. However, reliable camera pose estimation and scene reconstruction from such unconstrained videos…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Shuo Sun , Torsten Sattler , Malcolm Mielle , Achim J. Lilienthal , Martin Magnusson

This paper investigates an open research challenge of reconstructing high-quality, large 3D open scenes from images. It is observed existing methods have various limitations, such as requiring precise camera poses for input and dense…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Chong Cheng , Gaochao Song , Yiyang Yao , Qinzheng Zhou , Gangjian Zhang , Hao Wang

Prompt-driven scene synthesis allows users to generate complete 3D environments from textual descriptions. Current text-to-scene methods often struggle with complex geometries and object transformations, and tend to show weak adherence to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Frédéric Berdoz , Luca A. Lanzendörfer , Nick Tuninga , Roger Wattenhofer

Reinforcement learning (RL) has shown impressive success in exploring high-dimensional environments to learn complex tasks, but can often exhibit unsafe behaviors and require extensive environment interaction when exploration is…

Machine Learning · Computer Science 2021-09-22 Albert Wilcox , Ashwin Balakrishna , Brijen Thananjeyan , Joseph E. Gonzalez , Ken Goldberg

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Jiawei Yang , Jiahui Huang , Yuxiao Chen , Yan Wang , Boyi Li , Yurong You , Apoorva Sharma , Maximilian Igl , Peter Karkus , Danfei Xu , Boris Ivanovic , Yue Wang , Marco Pavone

We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a "foundation" model for monocular depth estimation and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Stanislaw Szymanowicz , Eldar Insafutdinov , Chuanxia Zheng , Dylan Campbell , João F. Henriques , Christian Rupprecht , Andrea Vedaldi