English

DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images

Computer Vision and Pattern Recognition 2025-09-29 v2

Abstract

We present DanceText, a training-free framework for multilingual text editing in images, designed to support complex geometric transformations and achieve seamless foreground-background integration. While diffusion-based generative models have shown promise in text-guided image synthesis, they often lack controllability and fail to preserve layout consistency under non-trivial manipulations such as rotation, translation, scaling, and warping. To address these limitations, DanceText introduces a layered editing strategy that separates text from the background, allowing geometric transformations to be performed in a modular and controllable manner. A depth-aware module is further proposed to align appearance and perspective between the transformed text and the reconstructed background, enhancing photorealism and spatial consistency. Importantly, DanceText adopts a fully training-free design by integrating pretrained modules, allowing flexible deployment without task-specific fine-tuning. Extensive experiments on the AnyWord-3M benchmark demonstrate that our method achieves superior performance in visual quality, especially under large-scale and complex transformation scenarios. Code is avaible at https://github.com/YuZhenyuLindy/DanceText.git.

Keywords

Cite

@article{arxiv.2504.14108,
  title  = {DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images},
  author = {Zhenyu Yu and Mohd Yamani Idna Idris and Hua Wang and Pei Wang and Rizwan Qureshi and Shaina Raza and Aman Chadha and Yong Xiang and Zhixiang Chen},
  journal= {arXiv preprint arXiv:2504.14108},
  year   = {2025}
}
R2 v1 2026-06-28T23:03:56.558Z