English
Related papers

Related papers: AnchorWeave: World-Consistent Video Generation wit…

200 papers

Dense 3D reconstruction from continuous image streams requires both accurate geometric aggregation and stable long-term memory management. Recent feed-forward reconstruction frameworks integrate observations through persistent memory…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Feifei Li , Qi Song , Chi Zhang , Rui Huang

Generative models have emerged as an essential building block for many image synthesis and editing tasks. Recent advances in this field have also enabled high-quality 3D or video content to be generated that exhibits either multi-view or…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Sherwin Bahmani , Jeong Joon Park , Despoina Paschalidou , Hao Tang , Gordon Wetzstein , Leonidas Guibas , Luc Van Gool , Radu Timofte

This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Jiacheng Chen , Ziyu Jiang , Mingfu Liang , Bingbing Zhuang , Jong-Chyi Su , Sparsh Garg , Ying Wu , Manmohan Chandraker

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

Prior panorama stitching approaches heavily rely on pairwise feature correspondences and are unable to leverage geometric consistency across multiple views. This leads to severe distortion and misalignment, especially in challenging scenes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Zhengdong Zhu , Weiyi Xue , Zuyuan Yang , Wenlve Zhou , Zhiheng Zhou

We introduce CinemaWorld, a generative augmented reality system that augments the viewer's physical surroundings with automatically generated mixed reality 3D content extracted from and synchronized with 2D movie scenes. Our system…

Human-Computer Interaction · Computer Science 2026-03-10 Keiichi Ihara , DaeHo Lee , Manato Abe , Hye-Young Jo , Ryo Suzuki

We present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Guojun Lei , Chi Wang , Hong Li , Rong Zhang , Yikai Wang , Weiwei Xu

High-resolution video generation has emerged as a crucial task in computer vision, with wide-ranging applications in entertainment, simulation, and data augmentation. However, generating temporally coherent and visually realistic videos…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Abhinav Sagar

Memory retention challenges in deep neural architectures have ongoing limitations in the ability to process and recall extended contextual information. Token dependencies degrade as sequence length increases, leading to a decline in…

Computation and Language · Computer Science 2025-03-26 Frederick Dillon , Gregor Halvorsen , Simon Tattershall , Magnus Rowntree , Gareth Vanderpool

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Ruihuang Li , Liyi Chen , Zhengqiang Zhang , Varun Jampani , Vishal M. Patel , Lei Zhang

Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Kirill Mazur , Gwangbin Bae , Andrew J. Davison

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar

Online free-viewpoint video (FVV) reconstruction is challenged by slow per-frame optimization, inconsistent motion estimation, and unsustainable storage demands. To address these challenges, we propose the Reconfigurable Continuum Gaussian…

Image and Video Processing · Electrical Eng. & Systems 2025-10-31 Jiaye Fu , Qiankun Gao , Chengxiang Wen , Yanmin Wu , Siwei Ma , Jiaqi Zhang , Jian Zhang

Single-image 3D generation has emerged as a prominent research topic, playing a vital role in virtual reality, 3D modeling, and digital content creation. However, existing methods face challenges such as a lack of multi-view geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Jinbo Yan , Alan Zhao , Yixin Hu

Unconditional scene inference and generation are challenging to learn jointly with a single compositional model. Despite encouraging progress on models that extract object-centric representations (''slots'') from images, unconditional…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Patrick Emami , Pan He , Sanjay Ranka , Anand Rangarajan

Text to video generation has emerged as a critical frontier in generative artificial intelligence, yet existing approaches struggle with maintaining temporal consistency, compositional understanding, and fine grained control over visual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Piyushkumar Patel

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Gene Chou , Charles Herrmann , Kyle Genova , Boyang Deng , Songyou Peng , Bharath Hariharan , Jason Y. Zhang , Noah Snavely , Philipp Henzler

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zhaoyang Li , Dongjun Qian , Kai Su , Qishuai Diao , Xiangyang Xia , Chang Liu , Wenfei Yang , Tianzhu Zhang , Zehuan Yuan

Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to executable actions.…

Machine Learning · Computer Science 2026-03-17 Christos Ziakas , Amir Bar , Alessandra Russo