InnoText: A Unified Model for Visual Text Generation and Editing
Abstract
Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing remains relatively underexplored. Unlike general image generation, visual text tasks demand precise structural regularity and legibility, which may pose additional challenges for small-scale text and non-Latin scripts such as Chinese. Existing UNet-based models often struggle to produce clear and coherent text, while DiT-based models, though more expressive, are typically limited to a single task, which may lead to redundant training pipelines, inconsistent visual styles, and reduced cross-task generalization. To address these challenges, we propose InnoText, a unified DiT-based framework capable of performing both text generation and editing within a single model. We introduce a Font Size-Aware Modulation (FSAM) module to enhance representations across font scales, a Small-Character Aware Augmentation strategy to improve fine-grained fidelity, and a Task-Specific Region Weighted Loss for adaptive optimization. To support training and evaluation, we also construct a high-quality bilingual (English-Chinese) visual text dataset covering diverse fonts, sizes, and backgrounds. Experimental results demonstrate that our method achieves superior generation accuracy and editing quality, producing visually appealing and realistic text images.
Cite
@article{arxiv.2607.22101,
title = {InnoText: A Unified Model for Visual Text Generation and Editing},
author = {Haowei Liu and Runze He and Jian Lu and Ao Ma and Run Ling and Ke Cao and Jiasong Feng and Wei Feng and Shuo Lu and Yexing Xu and Yun Wang and Jing Wang and Zhanjie Zhang},
journal= {arXiv preprint arXiv:2607.22101},
year = {2026}
}
Comments
Accepted by ECCV 2026