English

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas

Human-Computer Interaction 2025-08-12 v1

Abstract

Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap, we propose Canvas3D, an interactive system leveraging a 3D engine to enable precise spatial manipulation for image generation. Upon user prompt, Canvas3D automatically converts textual descriptions into interactive objects within a 3D engine-driven virtual canvas, empowering direct and precise spatial configuration. These user-defined arrangements generate explicit spatial constraints that guide generative models in accurately reflecting user intentions in the resulting images. We conducted a closed-end comparative study between Canvas3D and a baseline system. And an open-ended study to evaluate our system "in the wild". The result indicates that Canvas3D outperforms the baseline on spatial control, interactivity, and overall user experience.

Keywords

Cite

@article{arxiv.2508.07135,
  title  = {Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas},
  author = {Runlin Duan and Yuzhao Chen and Rahul Jain and Yichen Hu and Jingyu Shi and Karthik Ramani},
  journal= {arXiv preprint arXiv:2508.07135},
  year   = {2025}
}
R2 v1 2026-07-01T04:42:45.274Z