中文
相关论文

相关论文: Comp-Attn: Present-and-Align Attention for Composi…

200 篇论文

Text-to-Image (T2I) Diffusion Models have achieved remarkable performance in generating high quality images. However, enabling precise control of continuous attributes, especially multiple attributes simultaneously, in a new domain (e.g.,…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Wonwoong Cho , Yan-Ying Chen , Matthew Klenk , David I. Inouye , Yanxia Zhang

Diffusion Transformers (DiT) excel at image and video generation but face computational challenges due to the quadratic complexity of self-attention operators. We propose DiTFastAttn, a post-training compression method to alleviate the…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Zhihang Yuan , Hanling Zhang , Pu Lu , Xuefei Ning , Linfeng Zhang , Tianchen Zhao , Shengen Yan , Guohao Dai , Yu Wang

Retrieval-augmented generation improves the factual accuracy of Large Language Models (LLMs) by incorporating external context, but often suffers from irrelevant retrieved content that hinders effectiveness. Context compression addresses…

计算与语言 · 计算机科学 2025-09-23 Lvzhou Luo , Yixuan Cao , Ping Luo

Despite recent advancements in text-to-image models, achieving semantically accurate images in text-to-image diffusion models is a persistent challenge. While existing initial latent optimization methods have demonstrated impressive…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Aravindan Sundaram , Ujjayan Pal , Abhimanyu Chauhan , Aishwarya Agarwal , Srikrishna Karanam

Learning compositional representation is a key aspect of object-centric learning as it enables flexible systematic generalization and supports complex visual reasoning. However, most of the existing approaches rely on auto-encoding…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Whie Jung , Jaehoon Yoo , Sungjin Ahn , Seunghoon Hong

With the rapid development of generative models, Artificial Intelligence-Generated Contents (AIGC) have exponentially increased in daily lives. Among them, Text-to-Video (T2V) generation has received widespread attention. Though many T2V…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Tengchuan Kou , Xiaohong Liu , Zicheng Zhang , Chunyi Li , Haoning Wu , Xiongkuo Min , Guangtao Zhai , Ning Liu

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand. Good synchronization between…

计算机视觉与模式识别 · 计算机科学 2018-12-17 Naji Khosravan , Shervin Ardeshir , Rohit Puri

Despite significant advancements in text-to-image models for generating high-quality images, these methods still struggle to ensure the controllability of text prompts over images in the context of complex text prompts, especially when it…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Zhenyu Wang , Enze Xie , Aoxue Li , Zhongdao Wang , Xihui Liu , Zhenguo Li

Despite recent advances in text-to-image (T2I) models, they often fail to faithfully render all elements of complex prompts, frequently omitting or misrepresenting specific objects and attributes. Test-time optimization has emerged as a…

Existing subject-driven text-to-image generation models suffer from tedious fine-tuning steps and struggle to maintain both text-image alignment and subject fidelity. For generating compositional subjects, it often encounters problems such…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Shengyuan Liu , Bo Wang , Ye Ma , Te Yang , Xipeng Cao , Quan Chen , Han Li , Di Dong , Peng Jiang

The recent surge in video generation has shown the growing demand for high-quality video synthesis using large vision models. Existing video generation models are predominantly based on the video diffusion transformer (vDiT), however, they…

硬件体系结构 · 计算机科学 2025-11-18 Wenxuan Miao , Yulin Sun , Aiyue Chen , Jing Lin , Yiwu Yao , Yiming Gan , Jieru Zhao , Jingwen Leng , Mingyi Guo , Yu Feng

We present an effective method for fusing visual-and-language representations for several question answering tasks including visual question answering and visual entailment. In contrast to prior works that concatenate unimodal…

计算机视觉与模式识别 · 计算机科学 2022-12-06 Maxwell Mbabilla Aladago , AJ Piergiovanni

Controllability is a fundamental requirement in video synthesis, where accurate alignment with conditioning signals is essential. Existing classifier-free guidance methods typically achieve conditioning indirectly by modeling the joint…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Weiqi Li , Zehao Zhang , Liang Lin , Guangrun Wang

Text-guided image-to-video (I2V) generation aims to generate a coherent video that preserves the identity of the input image and semantically aligns with the input prompt. Existing methods typically augment pretrained text-to-video (T2V)…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Xun Guo , Mingwu Zheng , Liang Hou , Yuan Gao , Yufan Deng , Pengfei Wan , Di Zhang , Yufan Liu , Weiming Hu , Zhengjun Zha , Haibin Huang , Chongyang Ma

Textual-visual matching aims at measuring similarities between sentence descriptions and images. Most existing methods tackle this problem without effectively utilizing identity-level annotations. In this paper, we propose an identity-aware…

计算机视觉与模式识别 · 计算机科学 2017-08-08 Shuang Li , Tong Xiao , Hongsheng Li , Wei Yang , Xiaogang Wang

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Zhe Cao , Tao Wang , Jiaming Wang , Yanghai Wang , Yuanxing Zhang , Jialu Chen , Miao Deng , Jiahao Wang , Yubin Guo , Chenxi Liao , Yize Zhang , Zhaoxiang Zhang , Jiaheng Liu

The burgeoning field of generative artificial intelligence has fundamentally reshaped our approach to content creation, with Large Vision-Language Models (LVLMs) standing at its forefront. While current LVLMs have demonstrated impressive…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Spencer Ramsey , Jeffrey Lee , Amina Grant

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Yili Li , Gang Xiong , Gaopeng Gou , Xiangyan Qu , Jiamin Zhuang , Zhen Li , Junzheng Shi

Video generation remains a challenging task due to spatiotemporal complexity and the requirement of synthesizing diverse motions with temporal consistency. Previous works attempt to generate videos in arbitrary lengths either in an…

计算机视觉与模式识别 · 计算机科学 2023-04-07 Xiaoqian Shen , Xiang Li , Mohamed Elhoseiny

Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn audio embedding in a latent space with text embedding as…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Shentong Mo , Jing Shi , Yapeng Tian