中文
相关论文

相关论文: GS-DiT: Advancing Video Generation with Pseudo 4D …

200 篇论文

3D Gaussian Splatting (3DGS) has emerged as a powerful representation due to its efficiency and high-fidelity rendering. 3DGS training requires a known camera pose for each input view, typically obtained by Structure-from-Motion (SfM)…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zhen-Hui Dong , Sheng Ye , Yu-Hui Wen , Nannan Li , Yong-Jin Liu

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here, we instead focus on the underexplored text-to-4D setting and synthesize…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Huan Ling , Seung Wook Kim , Antonio Torralba , Sanja Fidler , Karsten Kreis

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr

Deformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a canonical GS representation. However, in filmmaking, tight budgets can result in sparse…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Adrian Azzarelli , Nantheera Anantrasirichai , David R Bull

3D Gaussian Splatting (3DGS) has demonstrated impressive performance in synthesizing novel views after training on a given set of viewpoints. However, its rendering quality deteriorates when the synthesized view deviates significantly from…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Jiatong Xia , Lingqiao Liu

We present GP-4DGS, a novel framework that integrates Gaussian Processes (GPs) into 4D Gaussian Splatting (4DGS) for principled probabilistic modeling of dynamic scenes. While existing 4DGS methods focus on deterministic reconstruction,…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Mijeong Kim , Jungtaek Kim , Bohyung Han

4D generation has made remarkable progress in synthesizing dynamic 3D objects from input text, images, or videos. However, existing methods often represent motion as an implicit deformation field, which limits direct control and…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Lifan Wu , Ruijie Zhu , Yubo Ai , Tianzhu Zhang

Despite recent successes in novel view synthesis using 3D Gaussian Splatting (3DGS), modeling scenes with sparse inputs remains a challenge. In this work, we address two critical yet overlooked issues in real-world sparse-input modeling:…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Yingji Zhong , Zhihao Li , Dave Zhenyu Chen , Lanqing Hong , Dan Xu

We consider the problem of novel-view synthesis (NVS) for dynamic scenes. Recent neural approaches have accomplished exceptional NVS results for static 3D scenes, but extensions to 4D time-varying scenes remain non-trivial. Prior efforts…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Yuanxing Duan , Fangyin Wei , Qiyu Dai , Yuhang He , Wenzheng Chen , Baoquan Chen

Generating dynamic 3D object from a single-view video is challenging due to the lack of 4D labeled data. An intuitive approach is to extend previous image-to-3D pipelines by transferring off-the-shelf image generation models such as score…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Zijie Pan , Zeyu Yang , Xiatian Zhu , Li Zhang

Sparse-view scene reconstruction often faces significant challenges due to the constraints imposed by limited observational data. These limitations result in incomplete information, leading to suboptimal reconstructions using existing…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Xiangyu Sun , Runnan Chen , Mingming Gong , Dong Xu , Tongliang Liu

Dynamic 4D Gaussian Splatting (4DGS) effectively extends the high-speed rendering capabilities of 3D Gaussian Splatting (3DGS) to represent volumetric videos. However, the large number of Gaussians, substantial temporal redundancies, and…

图形学 · 计算机科学 2026-01-14 Hyeongmin Lee , Kyungjune Baek

Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines are based on implicit generative models, which are hard to…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Wenhao Hu , Xuexiang Wen , Xi Li , Gaoang Wang

Recent advances in 3D content creation mostly leverage optimization-based 3D generation via score distillation sampling (SDS). Though promising results have been exhibited, these methods often suffer from slow per-sample optimization,…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Jiaxiang Tang , Jiawei Ren , Hang Zhou , Ziwei Liu , Gang Zeng

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet,…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Junpeng Jiang , Gangyi Hong , Miao Zhang , Hengtong Hu , Kun Zhan , Rui Shao , Liqiang Nie

3D Gaussian Splatting (3DGS) enables high-fidelity real-time rendering, a key requirement for immersive applications. However, the extension of 3DGS to dynamic scenes remains limitations on the substantial data volume of dense Gaussians and…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Jiayu Yang , Weijian Su , Songqian Zhang , Yuqi Han , Jinli Suo , Qiang Zhang

3D editing plays a crucial role in editing and reusing existing 3D assets, thereby enhancing productivity. Recently, 3DGS-based methods have gained increasing attention due to their efficient rendering and flexibility. However, achieving…

图形学 · 计算机科学 2024-12-12 Yian Zhao , Wanshi Xu , Yang Wu , Weiheng Huang , Zhongqian Sun , Wei Yang

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

3D Shape represented as point cloud has achieve advancements in multimodal pre-training to align image and language descriptions, which is curial to object identification, classification, and retrieval. However, the discrete representations…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Haoyuan Li , Yanpeng Zhou , Yihan Zeng , Hang Xu , Xiaodan Liang

Generating high-quality 3D content requires models capable of learning robust distributions of complex scenes and the real-world objects within them. Recent Gaussian-based 3D reconstruction techniques have achieved impressive results in…

图像与视频处理 · 电气工程与系统科学 2024-12-16 Kevin Miao , Harsh Agrawal , Qihang Zhang , Federico Semeraro , Marco Cavallo , Jiatao Gu , Alexander Toshev