English
Related papers

Related papers: 4DGen: Grounded 4D Content Generation with Spatial…

200 papers

Ground-truth RGBD data are fundamental for a wide range of computer vision applications; however, those labeled samples are difficult to collect and time-consuming to produce. A common solution to overcome this lack of data is to employ…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 L. Papa , P. Russo , I. Amerini

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Shaoheng Fang , Hanwen Jiang , Yunpeng Bai , Niloy J. Mitra , Qixing Huang

Volumetric video has emerged as a key medium for immersive telepresence and augmented/virtual reality, enabling six-degrees-of-freedom (6DoF) navigation and realistic spatial interactions. However, delivering high-quality dynamic volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Houqiang Zhong , Zihan Zheng , Qiang Hu , Yuan Tian , Ning Cao , Lan Xu , Xiaoyun Zhang , Zhengxue Cheng , Li Song , Wenjun Zhang

Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descriptions, depicting…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Mingyuan Zhang , Huirong Li , Zhongang Cai , Jiawei Ren , Lei Yang , Ziwei Liu

3D meshes are widely used in computer vision and graphics for their efficiency in animation and minimal memory use, playing a crucial role in movies, games, AR, and VR. However, creating temporally consistent and realistic textures for mesh…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Jingzhi Bao , Xueting Li , Ming-Hsuan Yang

Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in data scarcity: the most extensive 3D datasets encompass…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Antoine Mercier , Ramin Nakhli , Mahesh Reddy , Rajeev Yasarla , Hong Cai , Fatih Porikli , Guillaume Berger

We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Ruizhi Shao , Youxin Pang , Zerong Zheng , Jingxiang Sun , Yebin Liu

We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Liao Shen , Xingyi Li , Huiqiang Sun , Juewen Peng , Ke Xian , Zhiguo Cao , Guosheng Lin

Recent advances in text-to-video generation have harnessed the power of diffusion models to create visually compelling content conditioned on text prompts. However, they usually encounter high computational costs and often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Jiaxi Lv , Yi Huang , Mingfu Yan , Jiancheng Huang , Jianzhuang Liu , Yifan Liu , Yafei Wen , Xiaoxin Chen , Shifeng Chen

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Wanhua Li , Renping Zhou , Jiawei Zhou , Yingwei Song , Johannes Herter , Minghan Qin , Gao Huang , Hanspeter Pfister

Text-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Yukun Huang , Jianan Wang , Yukai Shi , Boshi Tang , Xianbiao Qi , Lei Zhang

Text-to-3D generation has shown great promise in generating novel 3D content based on given text prompts. However, existing generative methods mostly focus on geometric or visual plausibility while ignoring precise physics perception for…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Qingshan Xu , Jiao Liu , Melvin Wong , Caishun Chen , Yew-Soon Ong

Music-driven dance generation is a challenging task as it requires strict adherence to genre-specific choreography while ensuring physically realistic and precisely synchronized dance sequences with the music's beats and rhythm. Although…

Graphics · Computer Science 2026-04-21 Xinran Liu , Xu Dong , Shenbin Qian , Diptesh Kanojia , Wenwu Wang , Zhenhua Feng

This paper targets high-fidelity and real-time view synthesis of dynamic 3D scenes at 4K resolution. Recently, some methods on dynamic view synthesis have shown impressive rendering quality. However, their speed is still limited when…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhen Xu , Sida Peng , Haotong Lin , Guangzhao He , Jiaming Sun , Yujun Shen , Hujun Bao , Xiaowei Zhou

Recent advancements in text-to-3D generation improve the visual quality of Score Distillation Sampling (SDS) and its variants by directly connecting Consistency Distillation (CD) to score distillation. However, due to the imbalance between…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jiahao Zhu , Zixuan Chen , Guangcong Wang , Xiaohua Xie , Yi Zhou

Recent progress in 3D reconstruction has enabled realistic 3D models from dense image captures, yet challenges persist with sparse views, often leading to artifacts in unseen areas. Recent works leverage Video Diffusion Models (VDMs) to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Wenbin Teng , Gonglin Chen , Haiwei Chen , Yajie Zhao

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Andreas Blattmann , Robin Rombach , Huan Ling , Tim Dockhorn , Seung Wook Kim , Sanja Fidler , Karsten Kreis

Text-to-3D generation aims to create 3D assets from text-to-image diffusion models. However, existing methods face an inherent bottleneck in generation quality because the widely-used objectives such as Score Distillation Sampling (SDS)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Zixuan Chen , Ruijie Su , Jiahao Zhu , Lingxiao Yang , Jian-Huang Lai , Xiaohua Xie

We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Kuan Heng Lin , Zhizheng Liu , Pablo Salamanca , Yash Kant , Ryan Burgert , Yuancheng Xu , Koichi Namekata , Yiwei Zhao , Bolei Zhou , Micah Goldblum , Paul Debevec , Ning Yu
‹ Prev 1 8 9 10 Next ›