English
Related papers

Related papers: Adapting VACE for Real-Time Autoregressive Video D…

200 papers

Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Yuanzhi Wang , Yong Li , Xiaoya Zhang , Xin Liu , Anbo Dai , Antoni B. Chan , Zhen Cui

Autoregressive models excel in modeling sequential dependencies by enforcing causal constraints, yet they struggle to capture complex bidirectional patterns due to their unidirectional nature. In contrast, mask-based models leverage…

Computation and Language · Computer Science 2024-09-18 S. Rohollah Hosseyni , Ali Ahmad Rahmani , S. Jamal Seyedmohammadi , Sanaz Seyedin , Arash Mohammadi

Video (camera) trajectory editing aims to synthesize new videos that follow user-defined camera paths while preserving scene content and plausibly inpainting previously unseen regions, upgrading amateur footage into professionally styled…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zhihao Shi , Kejia Yin , Weilin Wan , Yuhongze Zhou , Yuanhao Yu , Xinxin Zuo , Qiang Sun , Juwei Lu

Variational Convertor-Encoder (VCE) converts an image to various styles; we present this novel architecture for the problem of one-shot generalization and its transfer to new tasks not seen before without additional training. We also…

Computer Vision and Pattern Recognition · Computer Science 2020-11-13 Chengshuai Li , Shuai Han , Jianping Xing

Generating high-quality novel views of a scene from a single image requires maintaining structural coherence across different views, referred to as view consistency. While diffusion models have driven advancements in novel view synthesis,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jiwoo Park , Tae Eun Choi , Youngjun Jun , Seong Jae Hwang

This paper proposes a non-autoregressive extension of our previously proposed sequence-to-sequence (S2S) model-based voice conversion (VC) methods. S2S model-based VC methods have attracted particular attention in recent years for their…

Sound · Computer Science 2021-04-15 Hirokazu Kameoka , Kou Tanaka , Takuhiro Kaneko

Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite rapid progress, most methods still require fixed-length inputs and substantial compute. Meanwhile,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Mohammadreza Salehi , Mehdi Noroozi , Luca Morreale , Ruchika Chavhan , Malcolm Chadwick , Alberto Gil Ramos , Abhinav Mehrotra

While latent diffusion models achieve impressive image editing results, their application to iterative editing of the same image is severely restricted. When trying to apply consecutive edit operations using current models, they accumulate…

Graphics · Computer Science 2025-04-29 Gal Almog , Ariel Shamir , Ohad Fried

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Tong Wang , Meng Zou , Chengjing Wu , Xiaochao Qu , Luoqi Liu , Xiaolin Hu , Ting Liu

Diffusion Transformer (DiT)-based video generation models inherently suffer from bottlenecks in long video synthesis and real-time inference, which can be attributed to the use of full spatiotemporal attention. Specifically, this mechanism…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Chao Yuan , Pan Li

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Yao-Chih Lee , Ji-Ze Genevieve Jang , Yi-Ting Chen , Elizabeth Qiu , Jia-Bin Huang

Autoregressive and diffusion models have achieved remarkable progress in language models and visual generation, respectively. We present ACDiT, a novel Autoregressive blockwise Conditional Diffusion Transformer, that innovatively combines…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Jinyi Hu , Shengding Hu , Yuxuan Song , Yufei Huang , Mingxuan Wang , Hao Zhou , Zhiyuan Liu , Wei-Ying Ma , Maosong Sun

Controllable generation, which enables fine-grained control over generated outputs, has emerged as a critical focus in visual generative models. Currently, there are two primary technical approaches in visual generation: diffusion models…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Ziyu Yao , Jialin Li , Yifeng Zhou , Yong Liu , Xi Jiang , Chengjie Wang , Feng Zheng , Yuexian Zou , Lei Li

Text-conditioned diffusion models have emerged as powerful tools for high-quality video generation. However, enabling Interactive Video Generation (IVG), where users control motion elements such as object trajectory, remains challenging.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Ishaan Rawal , Suryansh Kumar

Novel view synthesis from a single image has been a cornerstone problem for many Virtual Reality applications that provide immersive experiences. However, most existing techniques can only synthesize novel views within a limited range of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Hung-Yu Tseng , Qinbo Li , Changil Kim , Suhib Alsisan , Jia-Bin Huang , Johannes Kopf

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kesen Zhao , Jiaxin Shi , Beier Zhu , Junbao Zhou , Xiaolong Shen , Yuan Zhou , Qianru Sun , Hanwang Zhang

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Yixuan Ren , Yang Zhou , Jimei Yang , Jing Shi , Difan Liu , Feng Liu , Mingi Kwon , Abhinav Shrivastava

The remarkable generative capabilities of diffusion models have motivated extensive research in both image and video editing. Compared to video editing which faces additional challenges in the time dimension, image editing has witnessed the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Wenqi Ouyang , Yi Dong , Lei Yang , Jianlou Si , Xingang Pan

Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers become increasingly efficient, the latency bottleneck inevitably shifts to VAE decoders. To…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Lunjie Zhu , Yushi Huang , Xingtong Ge , Yufei Xue , Zhening Liu , Yumeng Zhang , Zehong Lin , Jun Zhang

Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuanhao Cai , He Zhang , Xi Chen , Jinbo Xing , Yiwei Hu , Yuqian Zhou , Kai Zhang , Zhifei Zhang , Soo Ye Kim , Tianyu Wang , Yulun Zhang , Xiaokang Yang , Zhe Lin , Alan Yuille
‹ Prev 1 4 5 6 7 8 10 Next ›