English
Related papers

Related papers: Explorative Inbetweening of Time and Space

200 papers

We study the problem of synthesizing immersive 3D indoor scenes from one or more images. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the input images while…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Jing Yu Koh , Harsh Agrawal , Dhruv Batra , Richard Tucker , Austin Waters , Honglak Lee , Yinfei Yang , Jason Baldridge , Peter Anderson

Visual generation has witnessed remarkable progress in single-image tasks, yet extending these capabilities to temporal sequences remains challenging. Current approaches either build specialized video models from scratch with enormous…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Cong Wan , Xiangyang Luo , Hao Luo , Zijian Cai , Yiren Song , Yunlong Zhao , Yifan Bai , Fan Wang , Yuhang He , Yihong Gong

Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 4D Gaussian Splatting (4DGS) have demonstrated impressive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zhixin Xu , Hengyu Zhou , Yuan Liu , Wenhan Xue , Hao Pan , Wenping Wang , Bin Wang

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Bing Li , Cheng Zheng , Wenxuan Zhu , Jinjie Mai , Biao Zhang , Peter Wonka , Bernard Ghanem

Modern video codecs and learning-based approaches struggle for semantic reconstruction at extremely low bit-rates due to reliance on low-level spatiotemporal redundancies. Generative models, especially diffusion models, offer a new paradigm…

Image and Video Processing · Electrical Eng. & Systems 2026-02-06 Maojun Zhang , Haotian Wu , Richeng Jin , Deniz Gunduz , Krystian Mikolajczyk

Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of each batch requires high cross-video diversity. Recent methods improve diversity for image…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Xinshuang Liu , Runfa Blark Li , Truong Nguyen

Generating new images with desired properties (e.g. new view/poses) from source images has been enthusiastically pursued recently, due to its wide range of potential applications. One way to ensure high-quality generation is to use multiple…

Computer Vision and Pattern Recognition · Computer Science 2022-02-03 Jiawei Lu , He Wang , Tianjia Shao , Yin Yang , Kun Zhou

Video inpainting aims to fill spatio-temporal holes with plausible content in a video. Despite tremendous progress of deep neural networks for image inpainting, it is challenging to extend these methods to the video domain due to the…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Dahun Kim , Sanghyun Woo , Joon-Young Lee , In So Kweon

Panoramic image stitching provides a unified, wide-angle view of a scene that extends beyond the camera's field of view. Stitching frames of a panning video into a panoramic photograph is a well-understood problem for stationary scenes, but…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Jingwei Ma , Erika Lu , Roni Paiss , Shiran Zada , Aleksander Holynski , Tali Dekel , Brian Curless , Michael Rubinstein , Forrester Cole

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Ting-Chun Wang , Ming-Yu Liu , Jun-Yan Zhu , Guilin Liu , Andrew Tao , Jan Kautz , Bryan Catanzaro

The application of video captioning models aims at translating the content of videos by using accurate natural language. Due to the complex nature inbetween object interaction in the video, the comprehensive understanding of spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Yutao Jin , Bin Liu , Jing Wang

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ruiyan Wang , Teng Hu , Kaihui Huang , Zihan Su , Ran Yi , Lizhuang Ma

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

The ability to generate complex and realistic human body animations at scale, while following specific artistic constraints, has been a fundamental goal for the game and animation industry for decades. Popular techniques include…

Computer Vision and Pattern Recognition · Computer Science 2020-05-29 Yi Zhou , Jingwan Lu , Connelly Barnes , Jimei Yang , Sitao Xiang , Hao li

Computer-assisted interventions can improve intra-operative guidance, particularly through deep learning methods that harness the spatiotemporal information in surgical videos. However, the severe data imbalance often found in surgical…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Danush Kumar Venkatesh , Isabel Funke , Micha Pfeiffer , Fiona Kolbinger , Hanna Maria Schmeiser , Juergen Weitz , Marius Distler , Stefanie Speidel

Coarse-guided visual generation, which synthesizes fine visual samples from degraded or low-fidelity coarse references, is essential for various real-world applications. While training-based approaches are effective, they are inherently…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yanghao Wang , Ziqi Jiang , Zhen Wang , Long Chen

Explorable 3D world generation from a single image or text prompt forms a cornerstone of spatial intelligence. Recent works utilize video model to achieve wide-scope and generalizable 3D world generation. However, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Zhongqi Yang , Wenhang Ge , Yuqi Li , Jiaqi Chen , Haoyuan Li , Mengyin An , Fei Kang , Hua Xue , Baixin Xu , Yuyang Yin , Eric Li , Yang Liu , Yikai Wang , Hao-Xiang Guo , Yahui Zhou

Event-based video reconstruction has garnered increasing attention due to its advantages, such as high dynamic range and rapid motion capture capabilities. However, current methods often prioritize the extraction of temporal information…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Lin Zhu , Yunlong Zheng , Yijun Zhang , Xiao Wang , Lizhi Wang , Hua Huang

We study the problem of synthesizing a number of likely future frames from a single input image. In contrast to traditional methods that have tackled this problem in a deterministic or non-parametric way, we propose to model future frames…

Computer Vision and Pattern Recognition · Computer Science 2019-08-13 Tianfan Xue , Jiajun Wu , Katherine L. Bouman , William T. Freeman

We present a feed-forward human performance capture method that renders novel views of a performer from a monocular RGB stream. A key challenge in this setting is the lack of sufficient observations, especially for unseen regions. Assuming…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Youngjoong Kwon , Yao He , Heejung Choi , Chen Geng , Zhengmao Liu , Jiajun Wu , Ehsan Adeli