English
Related papers

Related papers: Does Semantic Noise Initialization Transfer from I…

200 papers

Video colorization task has recently attracted wide attention. Recent methods mainly work on the temporal consistency in adjacent frames or frames with small interval. However, it still faces severe challenge of the inconsistency between…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Yu Zhang , Siqi Chen , Mingdao Wang , Xianlin Zhang , Chuang Zhu , Yue Zhang , Xueming Li

Recent advancements have integrated camera pose as a user-friendly and physics-informed condition in video diffusion models, enabling precise camera control. In this paper, we identify one of the key challenges as effectively modeling noisy…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Guangcong Zheng , Teng Li , Rui Jiang , Yehao Lu , Tao Wu , Xi Li

Semantic video segmentation is challenging due to the sheer amount of data that needs to be processed and labeled in order to construct accurate models. In this paper we present a deep, end-to-end trainable methodology to video segmentation…

Computer Vision and Pattern Recognition · Computer Science 2017-10-03 David Nilsson , Cristian Sminchisescu

Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelity, instruction-following content across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligned with text. Despite…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Kaisi Guan , Xihua Wang , Zhengfeng Lai , Xin Cheng , Peng Zhang , XiaoJiang Liu , Ruihua Song , Meng Cao

Visual generative AI models often encounter challenges related to text-image alignment and reasoning limitations. This paper presents a novel method for selectively enhancing the signal at critical denoising steps, optimizing image…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Paul Grimal , Hervé Le Borgne , Olivier Ferret

Text-to-video (T2V) synthesis models, such as OpenAI's Sora, have garnered significant attention due to their ability to generate high-quality videos from a text prompt. In diffusion-based T2V models, the attention mechanism is a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Bingyan Liu , Chengyu Wang , Tongtong Su , Huan Ten , Jun Huang , Kailing Guo , Kui Jia

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

Recent studies have made notable progress in video representation learning by transferring image-pretrained models to video tasks, typically with complex temporal modules and video fine-tuning. However, fine-tuning heavy modules may…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yang Liu , Qianqian Xu , Peisong Wen , Siran Dai , Xilin Zhao , Qingming Huang

Recent advances in text-to-video (T2V) diffusion models have significantly enhanced the quality of generated videos. However, their capability to produce explicit or harmful content introduces new challenges related to misuse and potential…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xiaoyu Ye , Songjie Cheng , Yongtao Wang , Yajiao Xiong , Yishen Li

We discover that common diffusion noise schedules do not enforce the last timestep to have zero signal-to-noise ratio (SNR), and some implementations of diffusion samplers do not start from the last timestep. Such designs are flawed and do…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Shanchuan Lin , Bingchen Liu , Jiashi Li , Xiao Yang

Concept erasure techniques for text-to-video (T2V) diffusion models report substantial suppression of sensitive content, yet current evaluation is limited to checking whether the target concept is absent from generated frames, treating…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yiwei Xie , Zheng Zhang , Ping Liu

Interactive video segmentation models such as SAM2 have demonstrated strong generalization across diverse visual domains. However, under weak user supervision, for example, when sparse point prompts are provided on a single frame, their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Dawar Jyoti Deka

Recent text-to-video (T2V) models can synthesize complex videos from lightweight natural language prompts, raising urgent concerns about safety alignment in the event of misuse in the real world. Prior jailbreak attacks typically rewrite…

Cryptography and Security · Computer Science 2026-03-10 Moyang Chen , Zonghao Ying , Wenzhuo Xu , Quancheng Zou , Deyue Zhang , Dongdong Yang , Xiangzheng Zhang

Text-to-image (T2I) diffusion models lack an efficient mechanism for early quality assessment, leading to costly trial-and-error in multi-generation scenarios such as prompt iteration, agent-based generation, and flow-grpo. We reveal a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Benlei Cui , Bukun Huang , Zhizeng Ye , Xuemei Dong , Tuo Chen , Hui Xue , Dingkang Yang , Longtao Huang , Jingqun Tang , Haiwen Hong

Many image-to-image (I2I) translation problems are in nature of high diversity that a single input may have various counterparts. Prior works proposed the multi-modal network that can build a many-to-many mapping between two visual domains.…

Computer Vision and Pattern Recognition · Computer Science 2019-10-07 Jialu Huang , Jing Liao , Tak Wu Sam Kwong

Text-to-Image (T2I) diffusion models enable high quality open ended synthesis, but practical use requires suppressing unsafe generations while preserving behavior on benign prompts. We study this tension relative to the frozen generator,…

Artificial Intelligence · Computer Science 2026-05-14 Minhyuk Lee , Hyekyung Yoon , Myungjoo Kang

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Kanghao Chen , Hangyu Li , JiaZhou Zhou , Zeyu Wang , Lin Wang

Instruction-guided generative models, especially those using text-to-image (T2I) and text-to-video (T2V) diffusion frameworks, have advanced the field of content editing in recent years. To extend these capabilities to 4D scene, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Hasan Iqbal , Nazmul Karim , Umar Khalid , Azib Farooq , Zichun Zhong , Chen Chen , Jing Hua