中文
相关论文

相关论文: Seeing World Dynamics in a Nutshell

200 篇论文

Recovering temporally consistent 3D human body pose, shape and motion from a monocular video is a challenging task due to (self-)occlusions, poor lighting conditions, complex articulated body poses, depth ambiguity, and limited availability…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Sushovan Chanda , Amogh Tiwari , Lokender Tiwari , Brojeshwar Bhowmick , Avinash Sharma , Hrishav Barua

High-fidelity reconstruction of 3D human avatars has a wild application in visual reality. In this paper, we introduce FAGhead, a method that enables fully controllable human portraits from monocular videos. We explicit the traditional 3D…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Yixin Xuan , Xinyang Li , Gongxin Yao , Shiwei Zhou , Donghui Sun , Xiaoxin Chen , Yu Pan

Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Shiqian Li , Ruihong Shen , Junfeng Ni , Chang Pan , Chi Zhang , Yixin Zhu

Recent advancements in dynamic 3D scene reconstruction have shown promising results, enabling high-fidelity 3D novel view synthesis with improved temporal consistency. Among these, 4D Gaussian Splatting (4DGS) has emerged as an appealing…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Seungjun Oh , Younggeun Lee , Hyejin Jeon , Eunbyung Park

Recent 4D reconstruction methods have yielded impressive results but rely on sharp videos as supervision. However, motion blur often occurs in videos due to camera shake and object movement, while existing methods render blurry results when…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Renlong Wu , Zhilu Zhang , Mingyang Chen , Zifei Yan , Wangmeng Zuo

Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing…

Reconstructing dynamic 3D scenes with photorealistic detail and strong temporal coherence remains a significant challenge. Existing Gaussian splatting approaches for dynamic scene modeling often rely on per-frame optimization, which can…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Tingxuan Huang , Haowei Zhu , Jun-hai Yong , Hao Pan , Bin Wang

The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its production remains costly and artifact-prone. To address this challenge, we present StereoWorld, an end-to-end framework that repurposes a…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Ke Xing , Xiaojie Jin , Longfei Li , Yuyang Yin , Hanwen Liang , Guixun Luo , Chen Fang , Jue Wang , Konstantinos N. Plataniotis , Yao Zhao , Yunchao Wei

The reconstruction of dynamic 3D scenes using 3D Gaussian Splatting has shown significant promise. A key challenge, however, remains in modeling realistic motion, as most methods fail to align the motion of Gaussians with real-world…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Junoh Lee , Junmyeong Lee , Yeon-Ji Song , Inhwan Bae , Jisu Shin , Hae-Gon Jeon , Jin-Hwa Kim

High-quality, animatable 3D human avatar reconstruction from monocular videos offers significant potential for reducing reliance on complex hardware, making it highly practical for applications in game development, augmented reality, and…

计算机视觉与模式识别 · 计算机科学 2025-05-02 Xia Yuan , Hai Yuan , Wenyi Ge , Ying Fu , Xi Wu , Guanyu Xing

Multi-object tracking (MOT) in monocular videos is fundamentally challenged by occlusions and depth ambiguity, issues that conventional tracking-by-detection (TBD) methods struggle to resolve owing to a lack of geometric awareness. To…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xudong Han , Pengcheng Fang , Yueying Tian , Jianhui Yu , Xiaohao Cai , Daniel Roggen , Philip Birch

Video-to-video synthesis (vid2vid) aims for converting high-level semantic inputs to photorealistic videos. While existing vid2vid methods can achieve short-term temporal consistency, they fail to ensure the long-term one. This is because…

计算机视觉与模式识别 · 计算机科学 2020-07-17 Arun Mallya , Ting-Chun Wang , Karan Sapra , Ming-Yu Liu

We consider the problem of novel-view synthesis (NVS) for dynamic scenes. Recent neural approaches have accomplished exceptional NVS results for static 3D scenes, but extensions to 4D time-varying scenes remain non-trivial. Prior efforts…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Yuanxing Duan , Fangyin Wei , Qiyu Dai , Yuhang He , Wenzheng Chen , Baoquan Chen

The field of 3D reconstruction from images has rapidly evolved in the past few years, first with the introduction of Neural Radiance Field (NeRF) and more recently with 3D Gaussian Splatting (3DGS). The latter provides a significant edge…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Avinash Paliwal , Wei Ye , Jinhui Xiong , Dmytro Kotovenko , Rakesh Ranjan , Vikas Chandra , Nima Khademi Kalantari

3D occupancy prediction is important for autonomous driving due to its comprehensive perception of the surroundings. To incorporate sequential inputs, most existing methods fuse representations from previous frames to infer the current 3D…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Sicheng Zuo , Wenzhao Zheng , Yuanhui Huang , Jie Zhou , Jiwen Lu

Realistic animatable human avatars from monocular videos are crucial for advancing human-robot interaction and enhancing immersive virtual experiences. While recent research on 3DGS-based human avatars has made progress, it still struggles…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Guangan Jiang , Tianzi Zhang , Dong Li , Zhenjun Zhao , Haoang Li , Mingrui Li , Hongyu Wang

Advances in Deep Learning have recently made it possible to recover full 3D meshes of human poses from individual images. However, extension of this notion to videos for recovering temporally coherent poses still remains unexplored. A major…

计算机视觉与模式识别 · 计算机科学 2019-07-02 Jian Liu , Naveed Akhtar , Ajmal Mian

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address this, we introduce Video-3DGS, a 3D Gaussian Splatting…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Inkyu Shin , Qihang Yu , Xiaohui Shen , In So Kweon , Kuk-Jin Yoon , Liang-Chieh Chen

Dynamic novel view synthesis aims to capture the temporal evolution of visual content within videos. Existing methods struggle to distinguishing between motion and structure, particularly in scenarios where camera poses are either unknown…

计算机视觉与模式识别 · 计算机科学 2024-01-12 Chaoyang Wang , Peiye Zhuang , Aliaksandr Siarohin , Junli Cao , Guocheng Qian , Hsin-Ying Lee , Sergey Tulyakov

While the generation of 3D content from single-view images has been extensively studied, the creation of physically consistent 3D dynamic scenes from videos remains in its early stages. We propose a novel framework leveraging generative 3D…

图形学 · 计算机科学 2025-10-31 Zhiwei Zhao , Alan Zhao , Minchen Li , Yixin Hu