English

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Computer Vision and Pattern Recognition 2026-02-11 v2

Abstract

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors predominantly support global or single-instruction edits and struggle with multi-instance scenarios, where multiple parts of a reference input must be edited independently without semantic interference. We identify this limitation as a consequence of globally conditioned velocity fields and joint attention mechanisms, which entangle concurrent edits. To address this issue, we introduce Instance-Disentangled Attention, a mechanism that partitions joint attention operations, enforcing binding between instance-specific textual instructions and spatial regions during velocity field estimation. We evaluate our approach on both natural image editing and a newly introduced benchmark of text-dense infographics with region-level editing instructions. Experimental results demonstrate that our approach promotes edit disentanglement and locality while preserving global output coherence, enabling single-pass, instance-level editing.

Keywords

Cite

@article{arxiv.2602.08749,
  title  = {Shifting the Breaking Point of Flow Matching for Multi-Instance Editing},
  author = {Carmine Zaccagnino and Fabio Quattrini and Enis Simsar and Marta Tintoré Gazulla and Rita Cucchiara and Alessio Tonioni and Silvia Cascianelli},
  journal= {arXiv preprint arXiv:2602.08749},
  year   = {2026}
}
R2 v1 2026-07-01T10:28:03.750Z