English
Related papers

Related papers: SpatialDreamer: Self-supervised Stereo Video Synth…

200 papers

Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties. While existing methods rely primarily on RGB features and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Jiaxin Cen , Xudong Mao , Guanghui Yue , Wei Zhou , Ruomei Wang , Fan Zhou , Baoquan Zhao

Diffusion models have demonstrated exceptional success in video super-resolution (VSR), exhibiting powerful capabilities for generating fine-grained details. However, their potential for space-time video super-resolution (STVSR), which…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Shuoyan Wei , Feng Li , Chen Zhou , Runmin Cong , Yao Zhao , Huihui Bai

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Wen-Hsuan Chu , Lei Ke , Katerina Fragkiadaki

In this paper, we firstly consider view-dependent effects into single image-based novel view synthesis (NVS) problems. For this, we propose to exploit the camera motion priors in NVS to model view-dependent appearance or effects (VDE) as…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Juan Luis Gonzalez Bello , Munchurl Kim

We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right view obtained by depth-based reprojection of the input left view. We extend the Stable Video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Nina Shvetsova , Goutam Bhat , Prune Truong , Hilde Kuehne , Federico Tombari

Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction remains an open…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Jiaxin Guo , Wenzhen Dong , Tianyu Huang , Hao Ding , Ziyi Wang , Haomin Kuang , Qi Dou , Yun-Hui Liu

Stereo vision is essential for many applications. Currently, the synchronization of the streams coming from two cameras is done using mostly hardware. A software-based synchronization method would reduce the cost, weight and size of the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Nicolas Boizard , Kevin El Haddad , Thierry Ravet , François Cresson , Thierry Dutoit

We present Stable Video 4D 2.0 (SV4D 2.0), a multi-view video diffusion model for dynamic 3D asset generation. Compared to its predecessor SV4D, SV4D 2.0 is more robust to occlusions and large motion, generalizes better to real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Chun-Han Yao , Yiming Xie , Vikram Voleti , Huaizu Jiang , Varun Jampani

Self-supervised Multi-view stereo (MVS) with a pretext task of image reconstruction has achieved significant progress recently. However, previous methods are built upon intuitions, lacking comprehensive explanations about the effectiveness…

Computer Vision and Pattern Recognition · Computer Science 2021-09-09 Hongbin Xu , Zhipeng Zhou , Yali Wang , Wenxiong Kang , Baigui Sun , Hao Li , Yu Qiao

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address this, we introduce Video-3DGS, a 3D Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Inkyu Shin , Qihang Yu , Xiaohui Shen , In So Kweon , Kuk-Jin Yoon , Liang-Chieh Chen

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junyoung Seo , Jisang Han , Jaewoo Jung , Siyoon Jin , Joungbin Lee , Takuya Narihira , Kazumi Fukuda , Takashi Shibuya , Donghoon Ahn , Shoukang Hu , Seungryong Kim , Yuki Mitsufuji

This paper tackles the challenges of self-supervised monocular depth estimation in indoor scenes caused by large rotation between frames and low texture. We ease the learning process by obtaining coarse camera poses from monocular sequences…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Chaoqiang Zhao , Matteo Poggi , Fabio Tosi , Lei Zhou , Qiyu Sun , Yang Tang , Stefano Mattoccia

Most existing algorithms for depth estimation from single monocular images need large quantities of metric groundtruth depths for supervised learning. We show that relative depth can be an informative cue for metric depth estimation and can…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Yuanzhouhan Cao , Tianqi Zhao , Ke Xian , Chunhua Shen , Zhiguo Cao , Shugong Xu

In the realm of text-to-3D generation, utilizing 2D diffusion models through score distillation sampling (SDS) frequently leads to issues such as blurred appearances and multi-faced geometry, primarily due to the intrinsically noisy nature…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Pengsheng Guo , Hans Hao , Adam Caccavale , Zhongzheng Ren , Edward Zhang , Qi Shan , Aditya Sankar , Alexander G. Schwing , Alex Colburn , Fangchang Ma

We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio). Such multi-view setups are prohibitively…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zihan Wang , Jeff Tan , Tarasha Khurana , Neehar Peri , Deva Ramanan

Most text-to-video (T2V) generators prioritize aesthetic quality, but often ignoring the spatial constraints in the generated videos. In this work, we present SPATIALALIGN, a self-improvement framework that enhances T2V models capabilities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Fengming Liu , Tat-Jen Cham , Chuanxia Zheng

We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Zhibing Li , Mengchen Zhang , Tong Wu , Jing Tan , Jiaqi Wang , Dahua Lin

Existing Subject-to-Video Generation (S2V) methods have achieved high-fidelity and subject-consistent video generation, yet remain constrained to single-view subject references. This limitation renders the S2V task reducible to an S2I + I2V…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Ziyang Song , Xinyu Gong , Bangya Liu , Zelin Zhao

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Onkar Susladkar , Jishu Sen Gupta , Chirag Sehgal , Sparsh Mittal , Rekha Singhal
‹ Prev 1 8 9 10 Next ›