HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction
Abstract
Drones operating in human-occupied spaces suffer from insufficient communication mechanisms that create uncertainty about their intentions. We present HoverAI, an embodied aerial agent that integrates drone mobility, infrastructure-independent visual projection, and real-time conversational AI into a unified platform. Equipped with a MEMS laser projector, onboard semi-rigid screen, and RGB camera, HoverAI perceives users through vision and voice, responding via lip-synced avatars that adapt appearance to user demographics. The system employs a multimodal pipeline combining VAD, ASR (Whisper), LLM-based intent classification, RAG for dialogue, face analysis for personalization, and voice synthesis (XTTS v2). Evaluation demonstrates high accuracy in command recognition (F1: 0.90), demographic estimation (gender F1: 0.89, age MAE: 5.14 years), and speech transcription (WER: 0.181). By uniting aerial robotics with adaptive conversational AI and self-contained visual output, HoverAI introduces a new class of spatially-aware, socially responsive embodied agents for applications in guidance, assistance, and human-centered interaction.
Cite
@article{arxiv.2601.13801,
title = {HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction},
author = {Yuhua Jin and Nikita Kuzmin and Georgii Demianchuk and Mariya Lezina and Fawad Mehboob and Issatay Tokmurziyev and Miguel Altamirano Cabrera and Muhammad Ahsan Mustafa and Dzmitry Tsetserukou},
journal= {arXiv preprint arXiv:2601.13801},
year = {2026}
}
Comments
This paper has been accepted for publication at LBR HRI 2026 conference