中文

AR 中的基于对象的叙事:一种结合视觉语言模型的情景-隐喻框架

人机交互 2025-04-18 v1

摘要

大多数自适应 AR 叙事系统通过简单对象标签和空间坐标来定义环境语义,限制叙事局限于刚性、预定义的逻辑。这种简化忽略了对象关系的情境意义——例如,夜灯上的婚戒可能暗示婚姻冲突,但仅被当作“两个空间中的对象”。为此,我们探索了将视觉语言模型(VLM)集成到 AR 流程中的方法。然而,several challenges 随之而来:First,stories generated with simple prompt guidance lacked narrative depth and spatial usage。Second,spatial semantics were underutilized,failing to support meaningful storytelling。Third,pre-generated scripts struggled to align with AR Foundation's object naming and coordinate systems。我们提出了一个 scene-driven AR 叙事框架,重新构想环境作为 active narrative agents,基于 three innovations:1. State-aware object semantics: We decompose object meaning into physical、functional、and metaphorical layers,allowing VLMs to distinguish subtle narrative cues between similar objects。2. Structured narrative interface: A bidirectional JSON layer maps VLM-generated metaphors to AR anchors,maintaining spatial and semantic coherence。3. STAM evaluation framework: A three-part experimental design evaluates narrative quality,highlighting both strengths and limitations of VLM-AR integration。我们的发现表明,该系统可以从环境本身生成故事,而不仅仅是将其置于环境之上。在用户研究中,70% 的参与者报告说,当叙事基于环境象征性时,会以不同的方式看待真实世界中的对象。通过将 VLMs 的生成创造力与 AR 的空间精度相结合,本框架引入了一种基于对象的叙事范式, 将被动空间转化为 active narrative landscapes。

关键词

引用

@article{arxiv.2504.13119,
  title  = {Object-Driven Narrative in AR: A Scenario-Metaphor Framework with VLM Integration},
  author = {Yusi Sun and Haoyan Guan and leith Kin Yep Chan and Yong Hong Kuo},
  journal= {arXiv preprint arXiv:2504.13119},
  year   = {2025}
}