English

PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation

Computer Vision and Pattern Recognition 2025-12-03 v1

Abstract

Recent diffusion-based text-to-image customization methods have achieved significant success in understanding concrete concepts to control generation processes, such as styles and shapes. However, few efforts dive into the realistic yet challenging customization of physical concepts. The core limitation of current methods arises from the absence of explicitly introducing physical knowledge during training. Even when physics-related words appear in the input text prompts, our experiments consistently demonstrate that these methods fail to accurately reflect the corresponding physical properties in the generated results. In this paper, we propose PhyCustom, a fine-tuning framework comprising two novel regularization losses to activate diffusion model to perform physical customization. Specifically, the proposed isometric loss aims at activating diffusion models to learn physical concepts while decouple loss helps to eliminate the mixture learning of independent concepts. Experiments are conducted on a diverse dataset and our benchmark results demonstrate that PhyCustom outperforms previous state-of-the-art and popular methods in terms of physical customization quantitatively and qualitatively.

Keywords

Cite

@article{arxiv.2512.02794,
  title  = {PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation},
  author = {Fan Wu and Cheng Chen and Zhoujie Fu and Jiacheng Wei and Yi Xu and Deheng Ye and Guosheng Lin},
  journal= {arXiv preprint arXiv:2512.02794},
  year   = {2025}
}

Comments

codes:https://github.com/wufan-cse/PhyCustom

R2 v1 2026-07-01T08:05:44.507Z