English
Related papers

Related papers: Anchor Token Matching: Implicit Structure Locking …

200 papers

The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, while autoregressive speech models have propelled…

Sound · Computer Science 2025-10-27 Yihan Wu , Georgios Milis , Ruibo Chen , Heng Huang

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to attention leakage and collision between the cross-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Xingxi Yin , Zhi Li , Jingfeng Zhang , Chenglin Li , Yin Zhang

Recent text-to-image diffusion models have significantly improved visual quality and text alignment. However, generating a sequence of images while preserving consistent character identity across diverse scene descriptions remains a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Shin Seong Kim , Minjung Shin , Hyunin Cho , Youngjung Uh

Text-conditioned image generation models are a prevalent use of AI image synthesis, yet intuitively controlling output guided by an artist remains challenging. Current methods require multiple images and textual prompts for each object to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Shounak Chatterjee

Consistent editing of real images is a challenging task, as it requires performing non-rigid edits (e.g., changing postures) to the main objects in the input image without changing their identity or attributes. To guarantee consistent…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Xiaoyue Duan , Shuhao Cui , Guoliang Kang , Baochang Zhang , Zhengcong Fei , Mingyuan Fan , Junshi Huang

Diffusion models have emerged as a powerful paradigm for modern generative modeling, demonstrating strong potential for large language models (LLMs). Unlike conventional autoregressive (AR) models that generate tokens sequentially,…

Machine Learning · Computer Science 2026-01-09 Gen Li , Changxiao Cai

Personalized image retouching aims to adapt retouching style of individual users from reference examples, but existing methods often require user-specific fine-tuning or fail to generalize effectively. To address these challenges, we…

Graphics · Computer Science 2026-02-20 Temesgen Muruts Weldengus , Binnan Liu , Fei Kou , Youwei Lyu , Jinwei Chen , Qingnan Fan , Changqing Zou

While diffusion models have achieved remarkable success in text-to-image generation, they encounter significant challenges with instruction-driven image editing. Our research highlights a key challenge: these models particularly struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Yujia Hu , Songhua Liu , Zhenxiong Tan , Xingyi Yang , Xinchao Wang

Text-to-video models have recently undergone rapid and substantial advancements. Nevertheless, due to limitations in data and computational resources, achieving efficient generation of long videos with rich motion dynamics remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Zongyi Li , Shujie Hu , Shujie Liu , Long Zhou , Jeongsoo Choi , Lingwei Meng , Xun Guo , Jinyu Li , Hefei Ling , Furu Wei

Generative text-to-image models are typically trained on large-scale web-scraped datasets that include diverse visual content such as copyrighted and stylistically distinctive artworks, raising concerns about ownership, attribution, and the…

Machine Learning · Computer Science 2026-05-19 Ninad Joshi , Ashutosh Ranjan , Vivek Srivastava , Shirish Karande

Existing 1D visual tokenizers for autoregressive (AR) generation largely follow the design principles of language modeling, as they are built directly upon transformers whose priors originate in language, yielding single-hierarchy latent…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Xu Zhang , Cheng Da , Huan Yang , Kun Gai , Ming Lu , Zhan Ma

Revolutionary advancements in text-to-image models have unlocked new dimensions for sophisticated content creation, such as text-conditioned image editing, enabling the modification of existing images based on textual guidance. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Haoyu Zheng , Wenqiao Zhang , Yaoke Wang , Juncheng Li , Zheqi Lv , Xin Min , Mengze Li , Dongping Zhang , Siliang Tang , Yueting Zhuang

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Recent advances in large language models (LLMs) have spurred interests in encoding images as discrete tokens and leveraging autoregressive (AR) frameworks for visual generation. However, the quantization process in AR-based visual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Peng Zheng , Junke Wang , Yi Chang , Yizhou Yu , Rui Ma , Zuxuan Wu

Training multimodal generative models on large, uncurated datasets can result in users being exposed to harmful, unsafe and controversial or culturally-inappropriate outputs. While model editing has been proposed to remove or filter…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Jordan Vice , Naveed Akhtar , Mubarak Shah , Richard Hartley , Ajmal Mian

Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats generated tokens as immutable, which may lead to error…

Diffusion models have recently achieved remarkable photorealism, making it increasingly difficult to distinguish real images from generated ones, raising significant privacy and security concerns. In response, we present a key finding:…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Wan Jiang , Jing Yan , Xiaojing Chen , Lin Shen , Chenhao Lin , Yunfeng Diao , Richang Hong

Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Zheng Gu , Shiyuan Yang , Jing Liao , Jing Huo , Yang Gao

Text-to-image synthesis has achieved high-quality results with recent advances in diffusion models. However, text input alone has high spatial ambiguity and limited user controllability. Most existing methods allow spatial control through…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yuki Endo

We present a training-free framework for continuous and controllable image editing at test time for text-conditioned generative models. In contrast to prior approaches that rely on additional training or manual user intervention, we find…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yigit Ekin , Yossi Gandelsman