English
Related papers

Related papers: CineScene: Implicit 3D as Effective Scene Represen…

200 papers

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Anyi Rao , Linning Xu , Yu Xiong , Guodong Xu , Qingqiu Huang , Bolei Zhou , Dahua Lin

We present a method that simultaneously addresses the tasks of dynamic scene novel-view synthesis and six degree-of-freedom (6-DOF) tracking of all dense scene elements. We follow an analysis-by-synthesis framework, inspired by recent work…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Jonathon Luiten , Georgios Kopanas , Bastian Leibe , Deva Ramanan

Our goal in this work is to generate realistic videos given just one initial frame as input. Existing unsupervised approaches to this task do not consider the fact that a video typically shows a 3D environment, and that this should remain…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Paul Henderson , Christoph H. Lampert , Bernd Bickel

While deep neural networks have led to human-level performance on computer vision tasks, they have yet to demonstrate similar gains for holistic scene understanding. In particular, 3D context has been shown to be an extremely important cue…

Computer Vision and Pattern Recognition · Computer Science 2017-08-17 Yinda Zhang , Mingru Bai , Pushmeet Kohli , Shahram Izadi , Jianxiong Xiao

A comprehensive semantic understanding of a scene is important for many applications - but in what space should diverse semantic information (e.g., objects, scene categories, material types, texture, etc.) be grounded and what should be its…

Computer Vision and Pattern Recognition · Computer Science 2019-10-08 Iro Armeni , Zhi-Yang He , JunYoung Gwak , Amir R. Zamir , Martin Fischer , Jitendra Malik , Silvio Savarese

3D understanding is a key capability for real-world AI assistance. High-quality data plays an important role in driving the development of the 3D understanding community. Current 3D scene understanding datasets often provide geometric and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zirui Wang , Tao Zhang

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

Careful robot manipulation in every-day cluttered environments requires an accurate understanding of the 3D scene, in order to grasp and place objects stably and reliably and to avoid colliding with other objects. In general, we must…

Robotics · Computer Science 2025-11-11 Aditya Agarwal , Gaurav Singh , Bipasha Sen , Tomás Lozano-Pérez , Leslie Pack Kaelbling

Controllable synthetic data generation can substantially lower the annotation cost of training data. Prior works use diffusion models to generate driving images conditioned on the 3D object layout. However, those models are trained on…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Yunsong Zhou , Michael Simon , Zhenghao Peng , Sicheng Mo , Hongzi Zhu , Minyi Guo , Bolei Zhou

We propose Text2Scene, a method to automatically create realistic textures for virtual scenes composed of multiple objects. Guided by a reference image and text descriptions, our pipeline adds detailed texture on labeled 3D geometries in…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Inwoo Hwang , Hyeonwoo Kim , Young Min Kim

We present BetterScene, an approach to enhance novel view synthesis (NVS) quality for diverse real-world scenes using extremely sparse, unconstrained photos. BetterScene leverages the production-ready Stable Video Diffusion (SVD) model…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Yuci Han , Charles Toth , John E. Anderson , William J. Shuart , Alper Yilmaz

Recent advancements in 3D object generation using diffusion models have achieved remarkable success, but generating realistic 3D urban scenes remains challenging. Existing methods relying solely on 3D diffusion models tend to suffer a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Hanlei Guo , Jiahao Shao , Xinya Chen , Xiyang Tan , Sheng Miao , Yujun Shen , Yiyi Liao

We propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model. We condition its UNet denoiser on the camera trajectory, using four…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Stefan Popov , Amit Raj , Michael Krainin , Yuanzhen Li , William T. Freeman , Michael Rubinstein

The entertainment industry relies on 3D visual content to create immersive experiences, but traditional methods for creating textured 3D models can be time-consuming and subjective. Generative networks such as StyleGAN have advanced image…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Yi-Ting Pan , Chai-Rong Lee , Shu-Ho Fan , Jheng-Wei Su , Jia-Bin Huang , Yung-Yu Chuang , Hung-Kuo Chu

Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, limiting the scene…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Lu Ling , Chen-Hsuan Lin , Tsung-Yi Lin , Yifan Ding , Yu Zeng , Yichen Sheng , Yunhao Ge , Ming-Yu Liu , Aniket Bera , Zhaoshuo Li

The visual world we sense, interpret and interact everyday is a complex composition of interleaved physical entities. Therefore, it is a very challenging task to generate vivid scenes of similar complexity using computers. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2019-02-05 Mehmet Ozgur Turkoglu , William Thong , Luuk Spreeuwers , Berkay Kicanaoglu

Incorporating camera intrinsics into video generation models offers a principled way to control not only scene dynamics but also the imaging process that governs visual appearance. Prior work has primarily focused on extrinsic control, such…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Debabrata Mandal , Zhihan Peng , Yujie Wang , Praneeth Chakravarthula

In the deep learning era, long video generation of high-quality still remains challenging due to the spatio-temporal complexity and continuity of videos. Existing prior works have attempted to model video distribution by representing videos…

Computer Vision and Pattern Recognition · Computer Science 2022-02-23 Sihyun Yu , Jihoon Tack , Sangwoo Mo , Hyunsu Kim , Junho Kim , Jung-Woo Ha , Jinwoo Shin

Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, complete,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Zeren Jiang , Yushi Lan , Yihang Luo , Yufan Deng , Zihang Lai , Edgar Sucar , Christian Rupprecht , Iro Laina , Diane Larlus , Chuanxia Zheng , Andrea Vedaldi

Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Gonzalo Gomez-Nogales , Yicong Hong , Chongjian Ge , Peiye Zhuang , Marc Comino-Trinidad , Dan Casas , Yi Zhou