English
Related papers

Related papers: Stabilizing Streaming Video Geometry via Dynamic F…

200 papers

Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 4D Gaussian Splatting (4DGS) have demonstrated impressive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zhixin Xu , Hengyu Zhou , Yuan Liu , Wenhan Xue , Hao Pan , Wenping Wang , Bin Wang

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Weixing Xie , Xiao Dong , Yong Yang , Qiqin Lin , Jingze Chen , Junfeng Yao , Xiaohu Guo

Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties. While existing methods rely primarily on RGB features and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Jiaxin Cen , Xudong Mao , Guanghui Yue , Wei Zhou , Ruomei Wang , Fan Zhou , Baoquan Zhao

We introduce MultiDiff, a novel approach for consistent novel view synthesis of scenes from a single RGB image. The task of synthesizing novel views from a single reference image is highly ill-posed by nature, as there exist multiple,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Norman Müller , Katja Schwarz , Barbara Roessle , Lorenzo Porzi , Samuel Rota Bulò , Matthias Nießner , Peter Kontschieder

We present an algorithm for estimating consistent dense depth maps and camera poses from a monocular video. We integrate a learning-based depth prior, in the form of a convolutional neural network trained for single-image depth estimation,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Johannes Kopf , Xuejian Rong , Jia-Bin Huang

Video stabilization often struggles with distortion and excessive cropping. This paper proposes a novel end-to-end framework, named TranStable, to address these challenges, comprising a genera tor and a discriminator. We establish…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 zhizhen li , tianyi zhuo , Yifei Cao , Jizhe Yu , Yu Liu

Existing 3D scene flow estimation methods provide the 3D geometry and 3D motion of a scene and gain a lot of interest, for example in the context of autonomous driving. These methods are traditionally based on a temporal series of stereo…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Fabian Brickwedde , Steffen Abraham , Rudolf Mester

The success of the GAN-NeRF structure has enabled face editing on NeRF to maintain 3D view consistency. However, achieving simultaneously multi-view consistency and temporal coherence while editing video sequences remains a formidable…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 Hao Zhang , Yu-Wing Tai , Chi-Keung Tang

While novel view synthesis (NVS) for dynamic scenes has seen significant progress, reconstructing temporally consistent geometric surfaces remains a challenge. Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) offer powerful…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Minje Kim , Younghyun Noh , Jaesoon Kim , Tae-Kyun Kim

We present an approach to generate a 360-degree view of a person with a consistent, high-resolution appearance from a single input image. NeRF and its variants typically require videos or images from different viewpoints. Most existing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Badour AlBahar , Shunsuke Saito , Hung-Yu Tseng , Changil Kim , Johannes Kopf , Jia-Bin Huang

In this paper, we tackle the problem of estimating the depth of a scene from a monocular video sequence. In particular, we handle challenging scenarios, such as non-translational camera motion and dynamic scenes, where traditional structure…

Computer Vision and Pattern Recognition · Computer Science 2015-11-20 Miaomiao Liu , Mathieu Salzmann , Xuming He

Visually exploring in a real-world 4D spatiotemporal space freely in VR has been a long-term quest. The task is especially appealing when only a few or even single RGB cameras are used for capturing the dynamic scene. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Liangchen Song , Anpei Chen , Zhong Li , Zhang Chen , Lele Chen , Junsong Yuan , Yi Xu , Andreas Geiger

Reconstructing scenes and tracking motion are two sides of the same coin. Tracking points allow for geometric reconstruction [14], while geometric reconstruction of (dynamic) scenes allows for 3D tracking of points over time [24, 39]. The…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Jenny Seidenschwarz , Qunjie Zhou , Bardienus Duisterhof , Deva Ramanan , Laura Leal-Taixé

Multi-frame methods improve monocular depth estimation over single-frame approaches by aggregating spatial-temporal information via feature matching. However, the spatial-temporal feature leads to accuracy degradation in dynamic scenes. To…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Jiquan Zhong , Xiaolin Huang , Xiao Yu

Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Dong Zhuo , Wenzhao Zheng , Jiahe Guo , Yuqi Wu , Jie Zhou , Jiwen Lu

In monocular videos that capture dynamic scenes, estimating the 3D geometry of video contents has been a fundamental challenge in computer vision. Specifically, the task is significantly challenged by the object motion, where existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Seong Hyeon Park , Jinwoo Shin

In recent years, advances in Artificial Intelligence have significantly impacted computer science, particularly in the field of computer vision, enabling solutions to complex problems such as video frame prediction. Video frame prediction…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jose M. Sánchez Velázquez , Mingbo Cai , Andrew Coney , Álvaro J. García- Tejedor , Alberto Nogales

Temporal quality is a critical aspect of video generation, as it ensures consistent motion and realistic dynamics across frames. However, achieving high temporal coherence and diversity remains challenging. In this work, we explore temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Harold Haodong Chen , Haojian Huang , Xianfeng Wu , Yexin Liu , Yajing Bai , Wen-Jie Shu , Harry Yang , Ser-Nam Lim

Monocular depth estimation is a crucial task in computer vision. While existing methods have shown impressive results under standard conditions, they often face challenges in reliably performing in scenarios such as low-light or rainy…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Yifan Mao , Jian Liu , Xianming Liu

While learning based depth estimation from images/videos has achieved substantial progress, there still exist intrinsic limitations. Supervised methods are limited by a small amount of ground truth or labeled data and unsupervised methods…

Computer Vision and Pattern Recognition · Computer Science 2019-05-24 Haofei Xu , Jianmin Zheng , Jianfei Cai , Juyong Zhang