English

Takeaways from Applying LLM Capabilities to Multiple Conversational Avatars in a VR Pilot Study

Human-Computer Interaction 2025-01-08 v2

Abstract

We present a virtual reality (VR) environment featuring conversational avatars powered by a locally-deployed LLM, integrated with automatic speech recognition (ASR), text-to-speech (TTS), and lip-syncing. Through a pilot study, we explored the effects of three types of avatar status indicators during response generation. Our findings reveal design considerations for improving responsiveness and realism in LLM-driven conversational systems. We also detail two system architectures: one using an LLM-based state machine to control avatar behavior and another integrating retrieval-augmented generation (RAG) for context-grounded responses. Together, these contributions offer practical insights to guide future work in developing task-oriented conversational AI in VR environments.

Keywords

Cite

@article{arxiv.2501.00168,
  title  = {Takeaways from Applying LLM Capabilities to Multiple Conversational Avatars in a VR Pilot Study},
  author = {Mykola Maslych and Christian Pumarada and Amirpouya Ghasemaghaei and Joseph J. LaViola},
  journal= {arXiv preprint arXiv:2501.00168},
  year   = {2025}
}
R2 v1 2026-06-28T20:52:55.979Z