English
Related papers

Related papers: PrevizWhiz: Combining Rough 3D Scenes and 2D Video…

200 papers

Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding,…

Machine Learning · Computer Science 2025-09-03 Xiaojie Xu , Xinli Xu , Sirui Chen , Haoyu Chen , Fan Zhang , Ying-Cong Chen

Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a challenging, resource-intensive process requiring deliberate…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Alejandro Pardo , Fabio Pizzati , Tong Zhang , Alexander Pondaven , Philip Torr , Juan Camilo Perez , Bernard Ghanem

We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short clips or rare transient…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Fufangchen Zhao , Liao Zhang , Daiqi Shi , Yuanjun Gao , Chen Ye , Yang Cai , Jian Gao , Danfeng Yan

Reconstructing a dynamic scene from image inputs is a fundamental computer vision task with many downstream applications. Despite recent advancements, existing approaches still struggle to achieve high-quality reconstructions from unseen…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Sara Oblak , Despoina Paschalidou , Sanja Fidler , Matan Atzmon

Video stylization, an important downstream task of video generation models, has not yet been thoroughly explored. Its input style conditions typically include text, style image, and stylized first frame. Each condition has a characteristic…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Mengtian Li , Jinshu Chen , Songtao Zhao , Wanquan Feng , Pengqi Tu , Qian He

We present a novel unconditional video generative model designed to address long-term spatial and temporal dependencies, with attention to computational and dataset efficiency. To capture long spatio-temporal dependencies, our approach…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Partha Ghosh , Soubhik Sanyal , Cordelia Schmid , Bernhard Schölkopf

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

Panoramic video generation enables immersive 360{\deg} content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Yifei Xia , Shuchen Weng , Siqi Yang , Jingqi Liu , Chengxuan Zhu , Minggui Teng , Zijian Jia , Han Jiang , Boxin Shi

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Recent advances in text-driven 3D scene editing and stylization, which leverage the powerful capabilities of 2D generative models, have demonstrated promising outcomes. However, challenges remain in ensuring high-quality stylization and…

Graphics · Computer Science 2026-03-03 Haruo Fujiwara , Yusuke Mukuta , Tatsuya Harada

Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens processed at each timestep. Recently, progressive…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Shikang Zheng , Jingkai Huang , Jiacheng Liu , Guantao Chen , Lixuan , Yuqi Lin , Peiliang Cai , Linfeng Zhang

Efficiently generating a freestyle 3D portrait with high quality and 3D-consistency is a promising yet challenging task. The portrait styles generated by most existing methods are usually restricted by their 3D generators, which are learned…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Tianxiang Ma , Kang Zhao , Jianxin Sun , Yingya Zhang , Jing Dong

Despite significant advances in modeling image priors via diffusion models, 3D-aware image editing remains challenging, in part because the object is only specified via a single image. To tackle this challenge, we propose 3D-Fixup, a new…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Yen-Chi Cheng , Krishna Kumar Singh , Jae Shin Yoon , Alex Schwing , Liangyan Gui , Matheus Gadelha , Paul Guerrero , Nanxuan Zhao

Recent successful video generation systems that predict and create realistic automotive driving scenes from short video inputs assign tokenization, future state prediction (world model), and video decoding to dedicated models. These…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Björn Möller , Zhengyang Li , Malte Stelzer , Thomas Graave , Fabian Bettels , Muaaz Ataya , Tim Fingscheidt

We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Zhiyuan Zhang , Dongdong Chen , Jing Liao

Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant progress has been made in generating 3D objects from text…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tianyu Huang , Wangguandong Zheng , Tengfei Wang , Yuhao Liu , Zhenwei Wang , Junta Wu , Jie Jiang , Hui Li , Rynson W. H. Lau , Wangmeng Zuo , Chunchao Guo

Estimating accurate camera poses, 3D scene geometry, and object motion from in-the-wild videos is a long-standing challenge for classical structure from motion pipelines due to the presence of dynamic objects. Recent learning-based methods…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Zhuoyuan Wu , Xurui Yang , Jiahui Huang , Yue Wang , Jun Gao

Generating photorealistic images with controllable camera pose and scene contents is essential for many applications including AR/VR and simulation. Despite the fact that rapid progress has been made in 3D-aware generative models, most…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Yuanbo Yang , Yifei Yang , Hanlei Guo , Rong Xiong , Yue Wang , Yiyi Liao

We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Rong Wang , Ruyi Zha , Ziang Cheng , Jiayu Yang , Pulak Purkait , Hongdong Li