English
Related papers

Related papers: Tracking-Guided 4D Generation: Foundation-Tracker …

200 papers

Existing multi-view 3D object reconstruction methods heavily rely on sufficient overlap between input views, where occlusions and sparse coverage in practice frequently yield severe reconstruction incompleteness. Recent advancements in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jiahao Chang , Chongjie Ye , Yushuang Wu , Yuantao Chen , Yidan Zhang , Zhongjin Luo , Chenghong Li , Yihao Zhi , Xiaoguang Han

Text-to-4D generation has recently been demonstrated viable by integrating a 2D image diffusion model with a video diffusion model. However, existing models tend to produce results with inconsistent motions and geometric structures over…

Graphics · Computer Science 2024-08-19 Ce Chen , Shaoli Huang , Xuelin Chen , Guangyi Chen , Xiaoguang Han , Kun Zhang , Mingming Gong

Generating dynamic 3D object from a single-view video is challenging due to the lack of 4D labeled data. An intuitive approach is to extend previous image-to-3D pipelines by transferring off-the-shelf image generation models such as score…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Zijie Pan , Zeyu Yang , Xiatian Zhu , Li Zhang

The field of generative models has recently witnessed significant progress, with diffusion models showing remarkable performance in image generation. In light of this success, there is a growing interest in exploring the application of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Ariel Lapid , Idan Achituve , Lior Bracha , Ethan Fetaya

Generative world models have become essential data engines for autonomous driving, yet most existing efforts focus on videos or occupancy grids, overlooking the unique LiDAR properties. Extending LiDAR generation to dynamic 4D world…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Ao Liang , Youquan Liu , Yu Yang , Dongyue Lu , Linfeng Li , Lingdong Kong , Huaici Zhao , Wei Tsang Ooi

Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Sicong Feng , Jielong Yang , Li Peng

Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 4D Gaussian Splatting (4DGS) have demonstrated impressive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zhixin Xu , Hengyu Zhou , Yuan Liu , Wenhan Xue , Hao Pan , Wenping Wang , Bin Wang

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Despite recent advances in leveraging generative prior from pre-trained diffusion models for 3D scene reconstruction, existing methods still face two critical limitations. First, due to the lack of reliable geometric supervision, they…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Junfeng Ni , Yixin Chen , Zhifei Yang , Yu Liu , Ruijie Lu , Song-Chun Zhu , Siyuan Huang

We train a feed-forward text-to-3D diffusion generator for human characters using only single-view 2D data for supervision. Existing 3D generative models cannot yet match the fidelity of image or video generative models. State-of-the-art 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Souhaib Attaiki , Paul Guerrero , Duygu Ceylan , Niloy J. Mitra , Maks Ovsjanikov

This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate videos at novel…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Yudong Jin , Sida Peng , Xuan Wang , Tao Xie , Zhen Xu , Yifan Yang , Yujun Shen , Hujun Bao , Xiaowei Zhou

Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Mingyang Wu , Ashirbad Mishra , Soumik Dey , Shuo Xing , Naveen Ravipati , Hansi Wu , Binbin Li , Zhengzhong Tu

Multi-view or 4D video generation has emerged as a significant research topic. Nonetheless, recent approaches to 4D generation still struggle with fundamental limitations, as they primarily rely on harnessing multiple video diffusion models…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Jangho Park , Taesung Kwon , Jong Chul Ye

Remarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Mengmeng Liu , Jiuming Liu , Yunpeng Zhang , Jiangtao Li , Michael Ying Yang , Francesco Nex , Hao Cheng

Recent advancements in generative models have enabled the creation of dynamic 4D content - 3D objects in motion - based on text prompts, which holds potential for applications in virtual worlds, media, and gaming. Existing methods provide…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Ohad Rahamim , Ori Malca , Dvir Samuel , Gal Chechik

Generating realistic 3D objects from single-view images requires natural appearance, 3D consistency, and the ability to capture multiple plausible interpretations of unseen regions. Existing approaches often rely on fine-tuning pretrained…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Pufan Li , Bi'an Du , Wei Hu

We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Rong Wang , Ruyi Zha , Ziang Cheng , Jiayu Yang , Pulak Purkait , Hongdong Li

Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Haoran Lu , Shang Wu , Jianshu Zhang , Maojiang Su , Guo Ye , Chenwei Xu , Lie Lu , Pranav Maneriker , Fan Du , Manling Li , Zhaoran Wang , Han Liu

The field of photorealistic 3D avatar reconstruction and generation has garnered significant attention in recent years; however, animating such avatars remains challenging. Recent advances in diffusion models have notably enhanced the…

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson