English

Text-to-Stage: Spatial Layouts from Long-form Narratives

Computation and Language 2026-03-19 v1 Artificial Intelligence Machine Learning

Abstract

In this work, we probe the ability of a language model to demonstrate spatial reasoning from unstructured text, mimicking human capabilities and automating a process that benefits many downstream media applications. Concretely, we study the narrative-to-play task: inferring stage-play layouts (scenes, speaker positions, movements, and room types) from text that lacks explicit spatial, positional, or relational cues. We then introduce a dramaturgy-inspired deterministic evaluation suite and, finally, a training and inference recipe that combines rejection SFT using Best-of-N sampling with RL from verifiable rewards via GRPO. Experiments on a text-only corpus of classical English literature demonstrate improvements over vanilla models across multiple metrics (character attribution, spatial plausibility, and movement economy), as well as alignment with an LLM-as-a-judge and subjective human preferences.

Keywords

Cite

@article{arxiv.2603.17832,
  title  = {Text-to-Stage: Spatial Layouts from Long-form Narratives},
  author = {Jefferson Hernandez and Swarnadeep Saha and Chenxi Whitehouse and Sanjeel Parekh and Calvin Murdock and Yuliang Li and W. Owen Brimijoin and Vamsi Krishna Ithapu and Ishwarya Ananthabhotla},
  journal= {arXiv preprint arXiv:2603.17832},
  year   = {2026}
}
R2 v1 2026-07-01T11:26:23.381Z