English

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

Computer Vision and Pattern Recognition 2026-05-20 v2

Abstract

Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on implicit generation processes that hinder precise control over scene structure and semantics. To address these limitations, we introduce RoomPilot, a unified framework for controllable indoor scene synthesis from multi-modal inputs, including textual descriptions and CAD floor plans. RoomPilot maps heterogeneous inputs into an Indoor Domain-Specific Language (IDSL), which serves as a structured and interpretable semantic representation for describing indoor scenes. Built upon IDSL, RoomPilot presents a hierarchical synthesis pipeline that progressively organizes scenes at the building, room, and object levels, promoting structural coherence and functional consistency across multi-room layouts. Moreover, RoomPilot constructs a curated asset dataset with rich semantic annotations to support high-quality scene synthesis, improving visual realism and appearance consistency. Extensive experiments demonstrate effective multi-modal understanding, fine-grained controllability in scene generation, and improved physical consistency and visual fidelity, marking a significant step toward controllable 3D indoor scene synthesis. Code and model will be available.

Keywords

Cite

@article{arxiv.2512.11234,
  title  = {RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing},
  author = {Wentang Chen and Shougao Zhang and Yiman Zhang and Tianhao Zhou and Ruihui Li},
  journal= {arXiv preprint arXiv:2512.11234},
  year   = {2026}
}

Comments

30 pages, 8 figures

R2 v1 2026-07-01T08:21:41.122Z