中文
相关论文

相关论文: Visual Prompting for One-shot Controllable Video E…

200 篇论文

Controllable ultra-long video generation is a fundamental yet challenging task. Although existing methods are effective for short clips, they struggle to scale due to issues such as temporal inconsistency and visual degradation. In this…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Jianxiong Gao , Zhaoxi Chen , Xian Liu , Jianfeng Feng , Chenyang Si , Yanwei Fu , Yu Qiao , Ziwei Liu

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Yao-Chih Lee , Ji-Ze Genevieve Jang , Yi-Ting Chen , Elizabeth Qiu , Jia-Bin Huang

Finding an initial noise vector that produces an input image when fed into the diffusion process (known as inversion) is an important problem in denoising diffusion models (DDMs), with applications for real image editing. The…

计算机视觉与模式识别 · 计算机科学 2022-12-23 Bram Wallace , Akash Gokul , Nikhil Naik

Text-based video editing has recently attracted considerable interest in changing the style or replacing the objects with a similar structure. Beyond this, we demonstrate that properties such as shape, size, location, motion, etc., can also…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Yue Ma , Xiaodong Cun , Sen Liang , Jinbo Xing , Yingqing He , Chenyang Qi , Siran Chen , Qifeng Chen

Text-conditional image editing is a very useful task that has recently emerged with immeasurable potential. Most current real image editing methods first need to complete the reconstruction of the image, and then editing is carried out by…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Songyan Chen , Jiancheng Huang

Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. Nevertheless, these methods suffer from severe content…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Ming Xie , Junqiu Yu , Qiaole Dong , Xiangyang Xue , Yanwei Fu

Leveraging pre-trained conditional diffusion models for video editing without further tuning has gained increasing attention due to its promise in film production, advertising, etc. Yet, seminal works in this line fall short in generation…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Zhenyi Liao , Zhijie Deng

Vision-centric autonomous driving systems rely on diverse and scalable training data to achieve robust performance. While video object editing offers a promising path for data augmentation, existing methods often struggle to maintain both…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Shuyun Wang , Haiyang Sun , Bing Wang , Hangjun Ye , Xin Yu

Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While deep learning-based methods have recently shown superior accuracy compared to traditional…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Dominik Kuczkowski , Laura Ruotsalainen

Diffusion models demonstrate impressive image generation performance with text guidance. Inspired by the learning process of diffusion, existing images can be edited according to text by DDIM inversion. However, the vanilla DDIM inversion…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Qi Qian , Haiyang Xu , Ming Yan , Juhua Hu

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often struggle with…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Jiangshan Wang , Junfu Pu , Zhongang Qi , Jiayi Guo , Yue Ma , Nisha Huang , Yuxin Chen , Xiu Li , Ying Shan

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Yangyang Xu , Shengfeng He , Kwan-Yee K. Wong , Ping Luo

The "one-shot" technique represents a distinct and sophisticated aesthetic in filmmaking. However, its practical realization is often hindered by prohibitive costs and complex real-world constraints. Although emerging video generation…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Jiawei Liu , Junqiao Li , Jiangfan Deng , Gen Li , Siyu Zhou , Zetao Fang , Shanshan Lao , Zengde Deng , Jianing Zhu , Tingting Ma , Jiayi Li , Yunqiu Wang , Qian He , Xinglong Wu

We introduce Videoshop, a training-free video editing algorithm for localized semantic edits. Videoshop allows users to use any editing software, including Photoshop and generative inpainting, to modify the first frame; it automatically…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Xiang Fan , Anand Bhattad , Ranjay Krishna

One image to editable dynamic 3D model and video generation is novel direction and change in the research area of single image to 3D representation or 3D reconstruction of image. Gaussian Splatting has demonstrated its advantages in…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Jinwei Lin

We present 3DScenePrompt, a framework that generates the next video chunk from arbitrary-length input while enabling precise camera control and preserving scene consistency. Unlike methods conditioned on a single image or a short clip, we…

计算机视觉与模式识别 · 计算机科学 2025-12-16 JoungBin Lee , Jaewoo Jung , Jisang Han , Takuya Narihira , Kazumi Fukuda , Junyoung Seo , Sunghwan Hong , Yuki Mitsufuji , Seungryong Kim

To get clear street-view and photo-realistic simulation in autonomous driving, we present an automatic video inpainting algorithm that can remove traffic agents from videos and synthesize missing regions with the guidance of depth/point…

计算机视觉与模式识别 · 计算机科学 2020-09-23 Miao Liao , Feixiang Lu , Dingfu Zhou , Sibo Zhang , Wei Li , Ruigang Yang

Our goal in this work is to generate realistic videos given just one initial frame as input. Existing unsupervised approaches to this task do not consider the fact that a video typically shows a 3D environment, and that this should remain…

计算机视觉与模式识别 · 计算机科学 2021-06-18 Paul Henderson , Christoph H. Lampert , Bernd Bickel

High-fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise video editing remains a formidable challenge. Modifying surgical attributes, such as…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Ritul Jangir , Arkya Jyoti Bagchi , Aiman Farooq , Mangalton Okram , Saurabh Seetaram Korgaonkar , Deepak Mishra

Even though large-scale text-to-image generative models show promising performance in synthesizing high-quality images, applying these models directly to image editing remains a significant challenge. This challenge is further amplified in…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Shutong Jin , Ruiyu Wang , Florian T. Pokorny