中文
相关论文

相关论文: Sync4D: Video Guided Controllable Dynamics for Phy…

200 篇论文

Recent advancements in generative models have enabled the creation of dynamic 4D content - 3D objects in motion - based on text prompts, which holds potential for applications in virtual worlds, media, and gaming. Existing methods provide…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Ohad Rahamim , Ori Malca , Dvir Samuel , Gal Chechik

Efficient neural representations for dynamic video scenes are critical for applications ranging from video compression to interactive simulations. Yet, existing methods often face challenges related to high memory usage, lengthy training…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Andrew Bond , Jui-Hsien Wang , Long Mai , Erkut Erdem , Aykut Erdem

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple image or video diffusion models, utilizing…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Hanwen Liang , Yuyang Yin , Dejia Xu , Hanxue Liang , Zhangyang Wang , Konstantinos N. Plataniotis , Yao Zhao , Yunchao Wei

Dynamic urban scene modeling is a rapidly evolving area with broad applications. While current approaches leveraging neural radiance fields or Gaussian Splatting have achieved fine-grained reconstruction and high-fidelity novel view…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yuru Xiao , Zihan Lin , Chao Lu , Deming Zhai , Kui Jiang , Wenbo Zhao , Wei Zhang , Junjun Jiang , Huanran Wang , Xianming Liu

Surgical scene simulation plays a crucial role in surgical education and simulator-based robot learning. Traditional approaches for creating these environments with surgical scene involve a labor-intensive process where designers hand-craft…

机器人学 · 计算机科学 2024-08-07 Zhenya Yang , Kai Chen , Yonghao Long , Qi Dou

High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm often suffers from…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Cheng Chi , Xianqi Wang , Hongcheng Luo , Mingfei Tu , Gangwei Xu , Zehan Zhang , Bing Wang , Guang Chen , Hangjun Ye , Sida Peng , Xin Yang , Haiyang Sun

Dynamic 3D reconstruction from monocular videos remains difficult due to the ambiguity inferring 3D motion from limited views and computational demands of modeling temporally varying scenes. While recent sparse control methods alleviate…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Jianing Chen , Zehao Li , Yujun Cai , Hao Jiang , Shuqin Gao , Honglong Zhao , Tianlu Mao , Yucheng Zhang

We address the challenge of generating 3D articulated objects in a controllable fashion. Currently, modeling articulated 3D objects is either achieved through laborious manual authoring, or using methods from prior work that are hard to…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Jiayi Liu , Hou In Ivan Tam , Ali Mahdavi-Amiri , Manolis Savva

Current 4D representations decouple geometry, motion, and semantics: reconstruction methods discard interpretable motion structure; language-grounded methods attach semantics after motion is learned, blind to how objects move; and…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Mohamed Rayan Barhdadi , Samir Abdaljalil , Rasul Khanbayov , Erchin Serpedin , Hasan Kurban

Recent 4D dynamic scene editing methods require editing thousands of 2D images used for dynamic scene synthesis and updating the entire scene with additional training loops, resulting in several hours of processing to edit a single dynamic…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Joohyun Kwon , Hanbyel Cho , Junmo Kim

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

Remarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Mengmeng Liu , Jiuming Liu , Yunpeng Zhang , Jiangtao Li , Michael Ying Yang , Francesco Nex , Hao Cheng

We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Zhibing Li , Mengchen Zhang , Tong Wu , Jing Tan , Jiaqi Wang , Dahua Lin

Capturing and re-animating the 3D structure of articulated objects present significant barriers. On one hand, methods requiring extensively calibrated multi-view setups are prohibitively complex and resource-intensive, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Heng Yu , Joel Julin , Zoltán Á. Milacski , Koichiro Niinuma , László A. Jeni

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr

3D Gaussian Splatting has shown fast and high-quality rendering results in static scenes by leveraging dense 3D prior and explicit representations. Unfortunately, the benefits of the prior and representation do not involve novel view…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Junoh Lee , Chang-Yeon Won , Hyunjun Jung , Inhwan Bae , Hae-Gon Jeon

3D Gaussian Splatting (3DGS) has substantial potential for enabling photorealistic Free-Viewpoint Video (FVV) experiences. However, the vast number of Gaussians and their associated attributes poses significant challenges for storage and…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Qiang Hu , Zihan Zheng , Houqiang Zhong , Sihua Fu , Li Song , XiaoyunZhang , Guangtao Zhai , Yanfeng Wang

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Chen Wang , Chuhao Chen , Yiming Huang , Zhiyang Dou , Yuan Liu , Jiatao Gu , Lingjie Liu

Controllable video generation has attracted significant attention, largely due to advances in video diffusion models. In domains such as autonomous driving, it is essential to develop highly accurate predictions for object motions. This…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Ge Ya Luo , Zhi Hao Luo , Anthony Gosselin , Alexia Jolicoeur-Martineau , Christopher Pal