English
Related papers

Related papers: Comp-Attn: Present-and-Align Attention for Composi…

200 papers

Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consistency throughout the video: existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Weiming Ren , Huan Yang , Ge Zhang , Cong Wei , Xinrun Du , Wenhao Huang , Wenhu Chen

In the realm of image synthesis, achieving fidelity to a reference image while adhering to conditional prompts remains a significant challenge. This paper proposes a novel approach that integrates a diffusion model with latent space…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Kshitij Pathania

In this paper, we focus on enhancing a diffusion-based text-to-video (T2V) model during the post-training phase by distilling a highly capable consistency model from a pretrained T2V model. Our proposed method, T2V-Turbo-v2, introduces a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Jiachen Li , Qian Long , Jian Zheng , Xiaofeng Gao , Robinson Piramuthu , Wenhu Chen , William Yang Wang

Recent works in dataset distillation seek to minimize training expenses by generating a condensed synthetic dataset that encapsulates the information present in a larger real dataset. These approaches ultimately aim to attain test accuracy…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Samir Khaki , Ahmad Sajedi , Kai Wang , Lucy Z. Liu , Yuri A. Lawryshyn , Konstantinos N. Plataniotis

Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Jiahao Wang , Caixia Yan , Weizhan Zhang , Haonan Lin , Mengmeng Wang , Guang Dai , Tieliang Gong , Hao Sun , Jingdong Wang

Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most recent works apply…

Computer Vision and Pattern Recognition · Computer Science 2024-03-04 Yuren Cong , Mengmeng Xu , Christian Simon , Shoufa Chen , Jiawei Ren , Yanping Xie , Juan-Manuel Perez-Rua , Bodo Rosenhahn , Tao Xiang , Sen He

Translating freehand sketches into photorealistic images remains a fundamental challenge in image synthesis, particularly due to the abstract, sparse, and stylistically diverse nature of sketches. Existing approaches, including GAN-based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Ali Zia , Muhammad Umer Ramzan , Usman Ali , Muhammad Faheem , Abdelwahed Khamis , Shahnawaz Qureshi

Recent text-to-video (T2V) models have demonstrated strong capabilities in producing high-quality, dynamic videos. To improve the visual controllability, recent works have considered fine-tuning pre-trained T2V models to support…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 June Suk Choi , Kyungmin Lee , Sihyun Yu , Yisol Choi , Jinwoo Shin , Kimin Lee

Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Kaisi Guan , Zhengfeng Lai , Yuchong Sun , Peng Zhang , Wei Liu , Kieran Liu , Meng Cao , Ruihua Song

The evolution of prompt learning methodologies has driven exploration of deeper prompt designs to enhance model performance. However, current deep text prompting approaches suffer from two critical limitations: Over-reliance on constrastive…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Qiqi Zhan , Shiwei Li , Qingjie Liu , Yunhong Wang

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a new…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yupeng Zhou , Daquan Zhou , Ming-Ming Cheng , Jiashi Feng , Qibin Hou

The task of Image-to-Video (I2V) generation aims to synthesize a video from a reference image and a text prompt. This requires diffusion models to reconcile high-frequency visual constraints and low-frequency textual guidance during the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yuanyang Yin , Yufan Deng , Shenghai Yuan , Kaipeng Zhang , Xiao Yang , Feng Zhao

Text-to-image (T2I) diffusion models generate high-quality images but often fail to capture the spatial relations specified in text prompts. This limitation can be traced to two factors: lack of fine-grained spatial supervision in training…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Sarah Rastegar , Violeta Chatalbasheva , Sieger Falkena , Anuj Singh , Yanbo Wang , Tejas Gokhale , Hamid Palangi , Hadi Jamali-Rad

Audio-visual video parsing (AVVP) aims to detect event categories and their temporal boundaries in videos, typically under weak supervision. Existing methods mainly focus on (i) improving temporal modeling using attention-based…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yaru Chen , Faegheh Sardari , Peiliang Zhang , Ruohao Guo , Yang Xiang , Zhenbo Li , Wenwu Wang

Temporal modeling is crucial for various video learning tasks. Most recent approaches employ either factorized (2D+1D) or joint (3D) spatial-temporal operations to extract temporal contexts from the input frames. While the former is more…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Yizhou Zhao , Zhenyang Li , Xun Guo , Yan Lu

Large-scale vision-language models (LVLMs) pretrained on massive image-text pairs have achieved remarkable success in visual representations. However, existing paradigms to transfer LVLMs to downstream tasks encounter two primary…

Multimedia · Computer Science 2023-08-01 Chunjin Yang , Fanman Meng , Shuai Chen , Mingyu Liu , Runtong Zhang

Compositional text-to-video generation, which requires synthesizing dynamic scenes with multiple interacting entities and precise spatial-temporal relationships, remains a critical challenge for diffusion-based models. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Weijie He , Mushui Liu , Yunlong Yu , Zhao Wang , Chao Wu

We present a lightweight latent diffusion model for vocal-conditioned musical accompaniment generation that addresses critical limitations in existing music AI systems. Our approach introduces a novel soft alignment attention mechanism that…

Sound · Computer Science 2026-01-06 Hei Shing Cheung , Boya Zhang , Jonathan H. Chan

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

Machine Learning · Computer Science 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Synthesizing images with user-specified subjects has received growing attention due to its practical applications. Despite the recent success in single subject customization, existing algorithms suffer from high training cost and low…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Zhiheng Liu , Yifei Zhang , Yujun Shen , Kecheng Zheng , Kai Zhu , Ruili Feng , Yu Liu , Deli Zhao , Jingren Zhou , Yang Cao
‹ Prev 1 3 4 5 6 7 10 Next ›