English

HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs

Robotics 2026-01-28 v1 Artificial Intelligence Human-Computer Interaction

Abstract

Existing human-robot interaction systems often lack mechanisms for sustained personalization and dynamic adaptation in multi-user environments, limiting their effectiveness in real-world deployments. We present HARMONI, a multimodal personalization framework that leverages large language models to enable socially assistive robots to manage long-term multi-user interactions. The framework integrates four key modules: (i) a perception module that identifies active speakers and extracts multimodal input; (ii) a world modeling module that maintains representations of the environment and short-term conversational context; (iii) a user modeling module that updates long-term speaker-specific profiles; and (iv) a generation module that produces contextually grounded and ethically informed responses. Through extensive evaluation and ablation studies on four datasets, as well as a real-world scenario-driven user-study in a nursing home environment, we demonstrate that HARMONI supports robust speaker identification, online memory updating, and ethically aligned personalization, outperforming baseline LLM-driven approaches in user modeling accuracy, personalization quality, and user satisfaction.

Keywords

Cite

@article{arxiv.2601.19839,
  title  = {HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs},
  author = {Jeanne Malécot and Hamed Rahimi and Jeanne Cattoni and Marie Samson and Mouad Abrini and Mahdi Khoramshahi and Maribel Pino and Mohamed Chetouani},
  journal= {arXiv preprint arXiv:2601.19839},
  year   = {2026}
}
R2 v1 2026-07-01T09:22:38.541Z