English
Related papers

Related papers: Stabilizing Streaming Video Geometry via Dynamic F…

200 papers

Why do deep networks generalize well? In contrast to classical generalization theory, we approach this fundamental question by examining not only inputs and outputs, but the evolution of internal features. Our study suggests a phenomenon of…

Machine Learning · Computer Science 2025-10-06 Tianyu Ruan , Kuo Gai , Shihua Zhang

The reconstruction of three-dimensional dynamic scenes is a well-established yet challenging task within the domain of computer vision. In this paper, we propose a novel approach that combines the domains of 3D geometry reconstruction and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 David Stotko , Reinhard Klein

Video monocular depth estimation is essential for applications such as autonomous driving, AR/VR, and robotics. Recent transformer-based single-image monocular depth estimation models perform well on single images but struggle with depth…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Sunghun Yang , Minhyeok Lee , Suhwan Cho , Jungho Lee , Sangyoun Lee

The performance of video prediction has been greatly boosted by advanced deep neural networks. However, most of the current methods suffer from large model sizes and require extra inputs, e.g., semantic/depth maps, for promising…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Xiaotao Hu , Zhewei Huang , Ailin Huang , Jun Xu , Shuchang Zhou

Dynamic radiance fields have emerged as a promising approach for generating novel views from a monocular video. However, previous methods enforce the geometric consistency to dynamic radiance fields only between adjacent input frames,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Byeongjun Park , Changick Kim

Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object deformation, texture drift, and non-rigid backgrounds under…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jan Ackermann , Shengqu Cai , Boyang Deng , Zhengfei Kuang , Songyou Peng , Gordon Wetzstein

Video frame prediction remains a fundamental challenge in computer vision with direct implications for autonomous systems, video compression, and media synthesis. We present FG-DFPN, a novel architecture that harnesses the synergy between…

Image and Video Processing · Electrical Eng. & Systems 2025-03-17 M. Akın Yılmaz , Ahmet Bilican , A. Murat Tekalp

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Hongyang Du , Junjie Ye , Xiaoyan Cong , Runhao Li , Jingcheng Ni , Aman Agarwal , Zeqi Zhou , Zekun Li , Randall Balestriero , Yue Wang

Convolutional neural networks (CNNs) can model complicated non-linear relations between images. However, they are notoriously sensitive to small changes in the input. Most CNNs trained to describe image-to-image mappings generate temporally…

Computer Vision and Pattern Recognition · Computer Science 2020-04-15 Gabriel Eilertsen , Rafał K. Mantiuk , Jonas Unger

Surface normal estimation serves as a cornerstone for a spectrum of computer vision applications. While numerous efforts have been devoted to static image scenarios, ensuring temporal coherence in video-based normal estimation remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Yanrui Bin , Wenbo Hu , Haoyuan Wang , Xinya Chen , Bing Wang

Recent advancements in dynamic neural radiance field methods have yielded remarkable outcomes. However, these approaches rely on the assumption of sharp input images. When faced with motion blur, existing dynamic NeRF methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Huiqiang Sun , Xingyi Li , Liao Shen , Xinyi Ye , Ke Xian , Zhiguo Cao

Neural Radiance Fields (NeRFs) excel in photorealistically rendering static scenes. However, rendering dynamic, long-duration radiance fields on ubiquitous devices remains challenging, due to data storage and computational constraints. In…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Liao Wang , Kaixin Yao , Chengcheng Guo , Zhirui Zhang , Qiang Hu , Jingyi Yu , Lan Xu , Minye Wu

Image stylization has seen significant advancement and widespread interest over the years, leading to the development of a multitude of techniques. Extending these stylization techniques, such as Neural Style Transfer (NST), to videos is…

Graphics · Computer Science 2023-07-03 Sumit Shekhar , Max Reimann , Moritz Hilscher , Amir Semmo , Jürgen Döllner , Matthias Trapp

Neural networks can represent and accurately reconstruct radiance fields for static 3D scenes (e.g., NeRF). Several works extend these to dynamic scenes captured with monocular video, with promising performance. However, the monocular…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Benjamin Attal , Eliot Laidlaw , Aaron Gokaslan , Changil Kim , Christian Richardt , James Tompkin , Matthew O'Toole

Online reconstruction of dynamic scenes is significant as it enables learning scenes from live-streaming video inputs, while existing offline dynamic reconstruction methods rely on recorded video inputs. However, previous online…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Youngsik Yun , Jeongmin Bae , Hyunseung Son , Seoha Kim , Hahyun Lee , Gun Bang , Youngjung Uh

Streaming feed-forward 3D reconstruction enables real-time joint estimation of scene geometry and camera poses from RGB images. However, without explicit dynamic reasoning, streaming models can be affected by moving objects, causing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Feiran Wang , Zezhou Shang , Gaowen Liu , Yan Yan

Recently, methods leveraging diffusion model priors to assist monocular geometric estimation (e.g., depth and normal) have gained significant attention due to their strong generalization ability. However, most existing works focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yang-Tian Sun , Xin Yu , Zehuan Huang , Yi-Hua Huang , Yuan-Chen Guo , Ziyi Yang , Yan-Pei Cao , Xiaojuan Qi

Dynamic scene reconstruction from monocular video is essential for real-world applications. We introduce DGNS, a hybrid framework integrating \underline{D}eformable \underline{G}aussian Splatting and Dynamic \underline{N}eural…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Xuesong Li , Jinguang Tong , Jie Hong , Vivien Rolland , Lars Petersson

The accurate reconstruction of dynamic scenes with neural radiance fields is significantly dependent on the estimation of camera poses. Widely used structure-from-motion pipelines encounter difficulties in accurately tracking the camera…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Nicolas Schischka , Hannah Schieber , Mert Asim Karaoglu , Melih Görgülü , Florian Grötzner , Alexander Ladikos , Daniel Roth , Nassir Navab , Benjamin Busam

Applying single image Monocular Depth Estimation (MDE) models to video sequences introduces significant temporal instability and flickering artifacts. We propose a novel approach that adapts any state-of-the-art image-based (depth)…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Ivan Sobko , Hayko Riemenschneider , Markus Gross , Christopher Schroers