English

OCTOPUS: Open-vocabulary Content Tracking and Object Placement Using Semantic Understanding in Mixed Reality

Computer Vision and Pattern Recognition 2023-12-21 v1 Artificial Intelligence Computation and Language

Abstract

One key challenge in augmented reality is the placement of virtual content in natural locations. Existing automated techniques are only able to work with a closed-vocabulary, fixed set of objects. In this paper, we introduce a new open-vocabulary method for object placement. Our eight-stage pipeline leverages recent advances in segmentation models, vision-language models, and LLMs to place any virtual object in any AR camera frame or scene. In a preliminary user study, we show that our method performs at least as well as human experts 57% of the time.

Keywords

Cite

@article{arxiv.2312.12815,
  title  = {OCTOPUS: Open-vocabulary Content Tracking and Object Placement Using Semantic Understanding in Mixed Reality},
  author = {Luke Yoffe and Aditya Sharma and Tobias Höllerer},
  journal= {arXiv preprint arXiv:2312.12815},
  year   = {2023}
}

Comments

IEEE International Symposium on Mixed and Augmented Reality (ISMAR) 2023