中文
相关论文

相关论文: MoVieS: Motion-Aware 4D Dynamic View Synthesis in …

200 篇论文

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Jiageng Mao , Boyi Li , Boris Ivanovic , Yuxiao Chen , Yan Wang , Yurong You , Chaowei Xiao , Danfei Xu , Marco Pavone , Yue Wang

This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising…

计算机视觉与模式识别 · 计算机科学 2025-03-28 David Yifan Yao , Albert J. Zhai , Shenlong Wang

We introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Hyeonho Jeong , Suhyeon Lee , Jong Chul Ye

In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseudo multi-view videos…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Yu Deng , Duomin Wang , Baoyuan Wang

Content creation, central to applications such as virtual reality, can be a tedious and time-consuming. Recent image synthesis methods simplify this task by offering tools to generate new views from as little as a single input image, or by…

计算机视觉与模式识别 · 计算机科学 2020-10-05 Tewodros Habtegebrial , Varun Jampani , Orazio Gallo , Didier Stricker

Gaussian Splatting (GS) has significantly elevated scene reconstruction efficiency and novel view synthesis (NVS) accuracy compared to Neural Radiance Fields (NeRF), particularly for dynamic scenes. However, current 4D NVS methods, whether…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Fang Li , Hao Zhang , Narendra Ahuja

Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Maximilian Seitzer , Sjoerd van Steenkiste , Thomas Kipf , Klaus Greff , Mehdi S. M. Sajjadi

We present a method for learning 3D geometry and physics parameters of a dynamic scene from only a monocular RGB video input. To decouple the learning of underlying scene geometry from dynamic motion, we represent the scene as a…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Yi-Ling Qiao , Alexander Gao , Ming C. Lin

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi

Synthesizing text-driven 3D human motion within realistic scenes requires learning both semantic intent ("walk to the couch") and physical feasibility (e.g., avoiding collisions). Current methods use generative frameworks that…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Anindita Ghosh , Vladislav Golyanik , Taku Komura , Philipp Slusallek , Christian Theobalt , Rishabh Dabral

We introduce Free3D, a simple accurate method for monocular open-set novel view synthesis (NVS). Similar to Zero-1-to-3, we start from a pre-trained 2D image generator for generalization, and fine-tune it for NVS. Compared to other works…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Chuanxia Zheng , Andrea Vedaldi

Rendering scenes observed in a monocular video from novel viewpoints is a challenging problem. For static scenes the community has studied both scene-specific optimization techniques, which optimize on every test scene, and generalized…

计算机视觉与模式识别 · 计算机科学 2024-02-21 Xiaoming Zhao , Alex Colburn , Fangchang Ma , Miguel Angel Bautista , Joshua M. Susskind , Alexander G. Schwing

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr

Motion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework,…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Wanquan Feng , Tianhao Qi , Jiawei Liu , Mingzhen Sun , Pengqi Tu , Tianxiang Ma , Fei Dai , Songtao Zhao , Siyu Zhou , Qian He

While recent years have witnessed great progress on using diffusion models for video generation, most of them are simple extensions of image generation frameworks, which fail to explicitly consider one of the key differences between videos…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Jingyun Liang , Yuchen Fan , Kai Zhang , Radu Timofte , Luc Van Gool , Rakesh Ranjan

We introduce a fully automatic pipeline for dynamic scene reconstruction from casually captured monocular RGB videos. Rather than designing a new scene representation, we enhance the priors that drive Dynamic Gaussian Splatting. Video…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Meng-Li Shih , Ying-Huan Chen , Yu-Lun Liu , Brian Curless

To achieve realistic immersion in landscape images, fluids such as water and clouds need to move within the image while revealing new scenes from various camera perspectives. Recently, a field called dynamic scene video has emerged, which…

计算机视觉与模式识别 · 计算机科学 2025-04-09 In-Hwan Jin , Haesoo Choo , Seong-Hun Jeong , Heemoon Park , Junghwan Kim , Oh-joon Kwon , Kyeongbo Kong

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi

Many motion-centric video analysis tasks, such as atomic actions, detecting atypical motor behavior in individuals with autism, or analyzing articulatory motion in real-time MRI of human speech, require efficient and interpretable temporal…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Hong Nguyen , Dung Tran , Hieu Hoang , Phong Nguyen , Shrikanth Narayanan

We present a novel neural radiance model that is trainable in a self-supervised manner for novel-view synthesis of dynamic unstructured scenes. Our end-to-end trainable algorithm learns highly complex, real-world static scenes within…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Shuja Khalid , Frank Rudzicz