English

SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences

Computer Vision and Pattern Recognition 2026-02-04 v1

Abstract

We introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D scene that reflects the real-world layout by compactly capturing the semantic cues of the surroundings. Prior works struggled to fully capture the contextual relationship between objects or mainly focused on synthesizing diverse shapes, making it challenging to generate 3D scenes aligned with object arrangements. We address these challenges by designing a graph network with cross-check feature attention for scene graph prediction and constructing a graph-variational autoencoder (graph-VAE), which consists of a joint shape and layout block for 3D scene generation. Experiments on the 3RScan/3DSSG and SG-FRONT datasets demonstrate that our approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations, even in complex indoor environments and under challenging scene graph constraints. Our work enables users to generate consistent 3D spaces from their physical environments via scene graphs, allowing them to create spatial MR content. Project page is https://scenelinker2026.github.io.

Keywords

Cite

@article{arxiv.2602.02974,
  title  = {SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences},
  author = {Seok-Young Kim and Dooyoung Kim and Woojin Cho and Hail Song and Suji Kang and Woontack Woo},
  journal= {arXiv preprint arXiv:2602.02974},
  year   = {2026}
}

Comments

Accepted as an IEEE TVCG paper at IEEE VR 2026 (journal track)

R2 v1 2026-07-01T09:33:16.848Z