English
Related papers

Related papers: Any4D: Unified Feed-Forward Metric 4D Reconstructi…

200 papers

Event cameras are rapidly emerging as powerful vision sensors for 3D reconstruction, uniquely capable of asynchronously capturing per-pixel brightness changes. Compared to traditional frame-based cameras, event cameras produce sparse yet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Chuanzhi Xu , Haoxian Zhou , Langyi Chen , Haodong Chen , Zeke Zexi Hu , Zhicheng Lu , Ying Zhou , Vera Chung , Qiang Qu , Weidong Cai

World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representations, lacking the comprehensive multi-view information essential for embodied spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Peiyan Tu , Hanxin Zhu , Jingwen Sun , Shaojie Ren , Cong Wang , Jiayi Luo , Xiaoqian Cheng , Zhibo Chen

Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datasets, these models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Kaichen Zhou , Yuhan Wang , Grace Chen , Xinhai Chang , Gaspard Beaudouin , Fangneng Zhan , Paul Pu Liang , Mengyu Wang

Large foundation models have recently emerged as a prominent focus of interest, attaining superior performance in widespread scenarios. Due to the scarcity of 3D data, many efforts have been made to adapt pre-trained transformers from…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Yiwen Tang , Ray Zhang , Jiaming Liu , Zoey Guo , Dong Wang , Zhigang Wang , Bin Zhao , Shanghang Zhang , Peng Gao , Hongsheng Li , Xuelong Li

Recovering 4D from monocular video, which jointly estimates dynamic geometry and camera poses, is an inevitably challenging problem. While recent pointmap-based 3D reconstruction methods (e.g., DUSt3R) have made great progress in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Shizun Wang , Zhenxiang Jiang , Xingyi Yang , Xinchao Wang

Fitting an underlying body model to 3D clothed human assets has been extensively studied, yet most approaches focus on either single-modal inputs such as point clouds or multi-view images alone, often requiring a known metric scale. This…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Zeyu Cai , Yuliang Xiu , Renke Wang , Zhijing Shao , Xiaoben Li , Siyuan Yu , Chao Xu , Yang Liu , Baigui Sun , Jian Yang , Zhenyu Zhang

We introduce Any6D, a model-free framework for 6D object pose estimation that requires only a single RGB-D anchor image to estimate both the 6D pose and size of unknown objects in novel scenes. Unlike existing methods that rely on textured…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Taeyeop Lee , Bowen Wen , Minjun Kang , Gyuree Kang , In So Kweon , Kuk-Jin Yoon

With the rapid development of 3D reconstruction technology, research in 4D reconstruction is also advancing, existing 4D reconstruction methods can generate high-quality 4D scenes. However, due to the challenges in acquiring multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Ling Yang , Kaixin Zhu , Juanxi Tian , Bohan Zeng , Mingbao Lin , Hongjuan Pei , Wentao Zhang , Shuicheng Yan

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Weixing Xie , Xiao Dong , Yong Yang , Qiqin Lin , Jingze Chen , Junfeng Yao , Xiaohu Guo

In this paper, we propose NeoVerse, a versatile 4D world model that is capable of 4D reconstruction, novel-trajectory video generation, and rich downstream applications. We first identify a common limitation of scalability in current 4D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yuxue Yang , Lue Fan , Ziqi Shi , Junran Peng , Feng Wang , Zhaoxiang Zhang

Humans have an innate ability to sense their surroundings, as they can extract the spatial representation from the egocentric perception and form an allocentric semantic map via spatial transformation and memory updating. However, endowing…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Chang Chen , Jiaming Zhang , Kailun Yang , Kunyu Peng , Rainer Stiefelhagen

We introduce Uni4D, a unified framework for large scale open vocabulary 3D retrieval and controlled 4D generation based on structured three level alignment across text, 3D models, and image modalities. Built upon the Align3D 130 dataset,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Philip Xu

We present Layout Anything, a transformer-based framework for indoor layout estimation that adapts the OneFormer's universal segmentation architecture to geometric structure prediction. Our approach integrates OneFormer's task-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Md Sohag Mia , Muhammad Abdullah Adnan

Existing diffusion models have made significant progress in generating realistic images. However, their direct adaptation to remote sensing imagery often disregards intrinsic physical laws. This oversight frequently leads to spectral…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zuopeng Zhao , Ying Liu , Xiaoyu Li , Su Luo , Lu Li , Wenwen Liu

The growing demand for rapid and scalable 3D asset creation has driven interest in feed-forward 3D reconstruction methods, with 3D Gaussian Splatting (3DGS) emerging as an effective scene representation. While recent approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Joanna Kaleta , Bartosz Świrta , Kacper Kania , Przemysław Spurek , Marek Kowalski

We propose a feed-forward Gaussian Splatting model that unifies 3D scene and semantic field reconstruction. Combining 3D scenes with semantic fields facilitates the perception and understanding of the surrounding environment. However, key…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Qijian Tian , Xin Tan , Jingyu Gong , Yuan Xie , Lizhuang Ma

Recent feed-forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi-view input, and unknown camera poses in a single…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Matteo Balice , Yanik Kunzi , Chenyangguang Zhang , Matteo Matteucci , Marc Pollefeys , Sungwhan Hong

We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a "foundation" model for monocular depth estimation and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Stanislaw Szymanowicz , Eldar Insafutdinov , Chuanxia Zheng , Dylan Campbell , João F. Henriques , Christian Rupprecht , Andrea Vedaldi

Volumetric video has emerged as a key medium for immersive telepresence and augmented/virtual reality, enabling six-degrees-of-freedom (6DoF) navigation and realistic spatial interactions. However, delivering high-quality dynamic volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Houqiang Zhong , Zihan Zheng , Qiang Hu , Yuan Tian , Ning Cao , Lan Xu , Xiaoyun Zhang , Zhengxue Cheng , Li Song , Wenjun Zhang

Reconstructing deformable surgical scenes from endoscopic videos is challenging and clinically important. Recent state-of-the-art methods based on implicit neural representations or 3D Gaussian splatting have made notable progress. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Jiwei Shan , Zeyu Cai , Cheng-Tai Hsieh , Yirui Li , Hao Liu , Lijun Han , Hesheng Wang , Shing Shin Cheng