English
Related papers

Related papers: InsViE-1M: Effective Instruction-based Video Editi…

200 papers

High-fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise video editing remains a formidable challenge. Modifying surgical attributes, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Ritul Jangir , Arkya Jyoti Bagchi , Aiman Farooq , Mangalton Okram , Saurabh Seetaram Korgaonkar , Deepak Mishra

In subject-driven text-to-image generation, recent works have achieved superior performance by training the model on synthetic datasets containing numerous image pairs. Trained on these datasets, generative models can produce text-aligned…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Yufan Zhou , Ruiyi Zhang , Kaizhi Zheng , Nanxuan Zhao , Jiuxiang Gu , Zichao Wang , Xin Eric Wang , Tong Sun

We introduce $\textit{InteractiveVideo}$, a user-centric framework for video generation. Different from traditional generative approaches that operate based on user-provided images or text, our framework is designed for dynamic interaction,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Yiyuan Zhang , Yuhao Kang , Zhixin Zhang , Xiaohan Ding , Sanyuan Zhao , Xiangyu Yue

We introduce a novel and efficient approach for text-based video-to-video editing that eliminates the need for resource-intensive per-video-per-model finetuning. At the core of our approach is a synthetic paired video dataset tailored for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Jiaxin Cheng , Tianjun Xiao , Tong He

We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with high temporal consistency, plausible scene transitions, and controllable streaming storylines. While existing long-video methods attempt to mitigate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Wuyang Li , Wentao Pan , Po-Chien Luan , Yang Gao , Alexandre Alahi

Vision-language large models have achieved remarkable success in various multi-modal tasks, yet applying them to video understanding remains challenging due to the inherent complexity and computational demands of video data. While…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Kai Han , Jianyuan Guo , Yehui Tang , Wei He , Enhua Wu , Yunhe Wang

Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from practical,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Cong Wei , Zheyang Xiong , Weiming Ren , Xinrun Du , Ge Zhang , Wenhu Chen

The rapid advancement in visual generation, particularly the emergence of pre-trained text-to-image and text-to-video models, has catalyzed growing interest in training-free video editing research. Mirroring training-free image editing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Lianghan Zhu , Yanqi Bao , Jing Huo , Jing Wu , Yu-Kun Lai , Wenbin Li , Yang Gao

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Juntong Wang , Jiarui Wang , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

Training vision-language models (VLMs) typically requires large-scale, high-quality image-text pairs, but collecting or synthesizing such data is costly. In contrast, text data is abundant and inexpensive, prompting the question: can…

Artificial Intelligence · Computer Science 2026-05-28 Xiaomin Yu , Wenjie Zhang , Ziyue Qiao , Chengwei Qin , Hui Xiong

The rapid growth of stereoscopic displays, including VR headsets and 3D cinemas, has led to increasing demand for high-quality stereo video content. However, producing 3D videos remains costly and complex, while automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Guibao Shen , Yihua Du , Wenhang Ge , Jing He , Chirui Chang , Donghao Zhou , Zhen Yang , Luozhou Wang , Xin Tao , Ying-Cong Chen

Despite significant advancements in video generation and editing using diffusion models, achieving accurate and localized video editing remains a substantial challenge. Additionally, most existing video editing methods primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Chong Mou , Mingdeng Cao , Xintao Wang , Zhaoyang Zhang , Ying Shan , Jian Zhang

Training-free video object editing aims to achieve precise object-level manipulation, including object insertion, swapping, and deletion. However, it faces significant challenges in maintaining fidelity and temporal consistency. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yiyang Chen , Xuanhua He , Xiujun Ma , Yue Ma

We introduce \textit{ImmersePro}, an innovative framework specifically designed to transform single-view videos into stereo videos. This framework utilizes a novel dual-branch architecture comprising a disparity branch and a context branch…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Jian Shi , Zhenyu Li , Peter Wonka

Though pre-training vision-language models have demonstrated significant benefits in boosting video-text retrieval performance from large-scale web videos, fine-tuning still plays a critical role with manually annotated clips with start and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Bin Zhu , Kevin Flanagan , Adriano Fragomeni , Michael Wray , Dima Damen

Vector format has been popular for representing icons and sketches. It has also been famous for design purposes. Regarding image editing, research on vector graphics editing rarely exists in contrast with the raster counterpart. We…

Graphics · Computer Science 2025-02-28 Kunato Nishina , Yusuke Matsui

The focus of this paper is on 3D motion editing. Given a 3D human motion and a textual description of the desired modification, our goal is to generate an edited motion as described by the text. The key challenges include the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Nikos Athanasiou , Alpár Cseke , Markos Diomataris , Michael J. Black , Gül Varol

Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using supervised vision-language instruction data. Recent studies have shown that large language models can…

Machine Learning · Computer Science 2026-04-14 Lai Wei , Xiaozhe Li , Zihao Jiang , Weiran Huang , Lichao Sun

Large-scale video generative models have recently demonstrated strong visual capabilities, enabling the prediction of future frames that adhere to the logical and physical cues in the current observation. In this work, we investigate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Gongfan Fang , Xinyin Ma , Xinchao Wang

Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Cheeun Hong , German Barquero , Fadime Sener , Markos Georgopoulos , Edgar Schönfeld , Stefan Popov , Yuming Du , Oscar Mañas , Albert Pumarola