LumiCtrl:用于个性化文本到图像模型中 lighting control 的 illuminant prompt 学习方法
计算机视觉与模式识别
2026-04-10 v2
摘要
文本到图像 (T2I) 模型在创意图像生成方面取得了显著进展,但仍缺乏对场景 illuminant 的精确控制,这是内容设计师操控生成图像视觉美学的关键因素。本文提出了一种名为 LumiCtrl 的illuminant 个性化方法,该方法根据单个物体图像学习 illuminant prompt。LumiCtrl 包含三个组件:给定物体图像,本方法应用 (a) 结合 Planckian locus 的 physics-based illuminant augmentation,以创建在标准 illuminant 下的微调变体;(b) 使用冻结的 ControlNet 进行 Edge-Guided Prompt Disentanglement,以确保 prompt 聚焦于光照而非结构;以及 (c) 一种 Masked Reconstruction Loss,专注于前景物体的学习,同时允许背景在语境中自适应,从而实现我们称之为 Contextual Light Adaptation。我们对 LumiCtrl 与其他 T2I 定制方法进行了定性和定量比较。结果表明,LumiCtrl 在 illuminant 保真度、审美质量和场景连贯性方面显著优于现有基线方法。人类偏好研究进一步确认了用户对 LumiCtrl 生成内容的强烈偏好。
引用
@article{arxiv.2512.17489,
title = {LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models},
author = {Muhammad Atif Butt and Kai Wang and Javier Vazquez-Corral and Joost Van De Weijer},
journal= {arXiv preprint arXiv:2512.17489},
year = {2026}
}
备注
Accepted to IEEE/CVF CVPR 2026 Workshop on AI for Creative Visual Content Generation, Editing, and Understanding (CVEU)