English
Related papers

Related papers: RunawayEvil: Jailbreaking the Image-to-Video Gener…

200 papers

In Image-to-Video (I2V) generation, a video is created using an input image as the first-frame condition. Existing I2V methods concatenate the full information of the conditional image with noisy latents to achieve high fidelity. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Yunyang Ge , Xinhua Cheng , Chengshu Zhao , Xianyi He , Shenghai Yuan , Bin Lin , Bin Zhu , Li Yuan

The rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic content, but it also raises significant security concerns. Current detection methods face two major…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Changjiang Jiang , Wenhui Dong , Zhonghao Zhang , Fengchang Yu , Wei Peng , Xinbin Yuan , Yifei Bi , Ming Zhao , Zian Zhou , Chenyang Si , Caifeng Shan

Diffusion-based image-to-video (I2V) models increasingly exhibit world-model-like properties by implicitly capturing temporal dynamics. However, existing studies have mainly focused on visual quality and controllability, and the robustness…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shuhan Xu , Siyuan Liang , Hongling Zheng , Yong Luo , Han Hu , Lefei Zhang , Dacheng Tao

We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual content generation. While existing models excel at generating…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Zhengjian Yao , Yongzhi Li , Xinyuan Gao , Quan Chen , Peng Jiang , Yanye Lu

The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Hengjia Li , Haonan Qiu , Shiwei Zhang , Xiang Wang , Yujie Wei , Zekun Li , Yingya Zhang , Boxi Wu , Deng Cai

The recent development of Sora leads to a new era in text-to-video (T2V) generation. Along with this comes the rising concern about its security risks. The generated videos may contain illegal or unethical content, and there is a lack of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yibo Miao , Yifan Zhu , Yinpeng Dong , Lijia Yu , Jun Zhu , Xiao-Shan Gao

The rapid rise of image-to-video (I2V) generation enables realistic videos to be created from a single image but also brings new forensic demands. Unlike static images, I2V content evolves over time, requiring forensics to move beyond 2D…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Yuzhuo Chen , Zehua Ma , Han Fang , Hengyi Wang , Guanjie Wang , Weiming Zhang

Recent progress in image generation models (IGMs) enables high-fidelity content creation but also amplifies risks, including the reproduction of copyrighted content and the generation of offensive content. Image Generation Model Unlearning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Yong Zou , Haoran Li , Fanxiao Li , Shenyang Wei , Yunyun Dong , Li Tang , Wei Zhou , Renyang Liu

As LLMs continue to shape real-world applications, automated jailbreak generation becomes essential to reveal safety weaknesses and guide model improvement. Existing automatic jailbreak generation methods have not yet fully considered two…

Neural and Evolutionary Computing · Computer Science 2026-05-06 Rui Tang , Kaiyu Xu , Pengsen Cheng , Hao Ren , Haizhou Wang , Shuyu Jiang

We introduce a training-free framework specifically designed to bring real-world static paintings to life through image-to-video (I2V) synthesis, addressing the persistent challenge of aligning these motions with textual guidance while…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Lingyu Liu , Yaxiong Wang , Li Zhu , Zhedong Zheng

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Jiasong Feng , Ao Ma , Jing Wang , Ke Cao , Zhanjie Zhang

While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization. We study the design space of data, architecture, and control, and introduce \emph{EasyV2V}, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jinjie Mai , Chaoyang Wang , Guocheng Gordon Qian , Willi Menapace , Sergey Tulyakov , Bernard Ghanem , Peter Wonka , Ashkan Mirzaei

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Multimodal learning involves developing models that can integrate information from various sources like images and texts. In this field, multimodal text generation is a crucial aspect that involves processing data from multiple modalities…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Youze Wang , Wenbo Hu , Richang Hong

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Ruicheng Zhang , Jun Zhou , Zunnan Xu , Zihao Liu , Jiehui Huang , Mingyang Zhang , Yu Sun , Xiu Li

Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially, text-to-video (T2V) generation, while many other works…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Ruibin Li , Tao Yang , Yangming Shi , Weiguo Feng , Shilei Wen , Bingyue Peng , Lei Zhang

Preserving first-frame identity while ensuring precise motion control is a fundamental challenge in human image animation. The Image-to-Motion Binding process of the dominant Reference-to-Video (R2V) paradigm overlooks critical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jiaming Zhang , Shengming Cao , Rui Li , Xiaotong Zhao , Yutao Cui , Xinglin Hou , Gangshan Wu , Haolan Chen , Yu Xu , Limin Wang , Kai Ma

Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models - a common and practical real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Linhao Huang , Xue Jiang , Zhiqiang Wang , Wentao Mo , Xi Xiao , Yong-Jie Yin , Bo Han , Feng Zheng

We introduce ``Idea to Image,'' a system that enables multimodal iterative self-refinement with GPT-4V(ision) for automatic image design and generation. Humans can quickly identify the characteristics of different text-to-image (T2I) models…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Zhengyuan Yang , Jianfeng Wang , Linjie Li , Kevin Lin , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

Recent AI systems have shown extremely powerful performance, even surpassing human performance, on various tasks such as information retrieval, language generation, and image generation based on large language models (LLMs). At the same…

Artificial Intelligence · Computer Science 2024-05-29 Minseon Kim , Hyomin Lee , Boqing Gong , Huishuai Zhang , Sung Ju Hwang
‹ Prev 1 3 4 5 6 7 10 Next ›