English
Related papers

Related papers: RealCam: Real-Time Novel-View Video Generation wit…

200 papers

Autoregressive video synthesis offers a promising pathway for infinite-horizon generation but is fundamentally hindered by three intertwined challenges: semantic forgetting from context limitations, visual drift due to positional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jintao Chen , Chengyu Bai , Junjun Hu , Xinda Xue , Mu Xu

Image-to-video (I2V) generation tasks always suffer from keeping high fidelity in the open domains. Traditional image animation techniques primarily focus on specific domains such as faces or human poses, making them difficult to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Weijie Li , Litong Gong , Yiran Zhu , Fanda Fan , Biao Wang , Tiezheng Ge , Bo Zheng

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yuancheng Xu , Wenqi Xian , Li Ma , Julien Philip , Ahmet Levent Taşel , Yiwei Zhao , Ryan Burgert , Mingming He , Oliver Hermann , Oliver Pilarski , Rahul Garg , Paul Debevec , Ning Yu

The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Kaichen Zhou , Jia-Xing Zhong , Sangyun Shin , Kai Lu , Yiyuan Yang , Andrew Markham , Niki Trigoni

Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges when dealing with complex scenes that require a deep…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Peng Liu , Xiaoming Ren , Fengkai Liu , Qingsong Xie , Quanlong Zheng , Yanhao Zhang , Haonan Lu , Yujiu Yang

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Ting-Chun Wang , Ming-Yu Liu , Jun-Yan Zhu , Guilin Liu , Andrew Tao , Jan Kautz , Bryan Catanzaro

With the rise of short video platforms represented by TikTok, the trend of users expressing their creativity through photos and videos has increased dramatically. However, ordinary users lack the professional skills to produce high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Sixiao Zheng , Jingyang Huo , Yu Wang , Yanwei Fu

High-quality driving video generation is crucial for providing training data for autonomous driving models. However, current generative models rarely focus on enhancing camera motion control under multi-view tasks, which is essential for…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Yining Yao , Xi Guo , Chenjing Ding , Wei Wu

Currently almost all state-of-the-art novel view synthesis and reconstruction models rely on calibrated cameras or additional geometric priors for training. These prerequisites significantly limit their applicability to massive uncalibrated…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Ruoyu Wang , Yi Ma , Shenghua Gao

Controlling video and audio generation requires diverse modalities, from depth and pose to camera trajectories and audio transformations, yet existing approaches either train a single monolithic model for a fixed set of controls or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Matan Ben-Yosef , Tavi Halperin , Naomi Ken Korem , Mohammad Salama , Harel Cain , Asaf Joseph , Anthony Chen , Urska Jelercic , Ofir Bibi

Rendering photo-realistic novel-view images of complex scenes has been a long-standing challenge in computer graphics. In recent years, great research progress has been made on enhancing rendering quality and accelerating rendering speed in…

Graphics · Computer Science 2024-02-21 Tiansong Zhou , Yebin Liu , Xuangeng Chu , Chengkun Cao , Changyin Zhou , Fei Yu , Yu Li

The narrative quality of a video fundamentally determines its perceptual value. Although existing video generation methods can produce visually appealing content, they predominantly rely on sparse conditioning signals such as text prompts…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Zhida Zhang , Jie Ma , Zhan Peng , Haoxue Wu , Yang Han , Jun Liang , Jie Cao , Jing Li

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there are currently no…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Shaoteng Liu , Yuechen Zhang , Wenbo Li , Zhe Lin , Jiaya Jia

Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, we introduce a large-scale, dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Zheng-Hui Huang , Zhixiang Wang , Jiaming Tan , Ruihan Yu , Yidan Zhang , Bo Zheng , Yu-Lun Liu , Yung-Yu Chuang , Kaipeng Zhang

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

VQ-based image generation typically follows a two-stage pipeline: a tokenizer encodes images into discrete tokens, and a generative model learns their dependencies for reconstruction. However, improved tokenization in the first stage does…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Bin Wu , Mengqi Huang , Weinan Jia , Zhendong Mao

Camera virtualization -- an emerging solution to novel view synthesis -- holds transformative potential for visual entertainment, live performances, and sports broadcasting by enabling the generation of photorealistic images from novel…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Yunxiao Zhang , William Stone , Suryansh Kumar

We propose 4Real-Video, a novel framework for generating 4D videos, organized as a grid of video frames with both time and viewpoint axes. In this grid, each row contains frames sharing the same timestep, while each column contains frames…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Chaoyang Wang , Peiye Zhuang , Tuan Duc Ngo , Willi Menapace , Aliaksandr Siarohin , Michael Vasilkovsky , Ivan Skorokhodov , Sergey Tulyakov , Peter Wonka , Hsin-Ying Lee

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

Artificial Intelligence · Computer Science 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

Emotional Video Captioning is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectively perceive subtle and ambiguous visual emotional cues during…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Cheng Ye , Weidong Chen , Jingyu Li , Lei Zhang , Zhendong Mao
‹ Prev 1 8 9 10 Next ›