English

Efficient Multimodal Neural Networks for Trigger-less Voice Assistants

Machine Learning 2023-05-23 v1 Human-Computer Interaction

Abstract

The adoption of multimodal interactions by Voice Assistants (VAs) is growing rapidly to enhance human-computer interactions. Smartwatches have now incorporated trigger-less methods of invoking VAs, such as Raise To Speak (RTS), where the user raises their watch and speaks to VAs without an explicit trigger. Current state-of-the-art RTS systems rely on heuristics and engineered Finite State Machines to fuse gesture and audio data for multimodal decision-making. However, these methods have limitations, including limited adaptability, scalability, and induced human biases. In this work, we propose a neural network based audio-gesture multimodal fusion system that (1) Better understands temporal correlation between audio and gesture data, leading to precise invocations (2) Generalizes to a wide range of environments and scenarios (3) Is lightweight and deployable on low-power devices, such as smartwatches, with quick launch times (4) Improves productivity in asset development processes.

Keywords

Cite

@article{arxiv.2305.12063,
  title  = {Efficient Multimodal Neural Networks for Trigger-less Voice Assistants},
  author = {Sai Srujana Buddi and Utkarsh Oggy Sarawgi and Tashweena Heeramun and Karan Sawnhey and Ed Yanosik and Saravana Rathinam and Saurabh Adya},
  journal= {arXiv preprint arXiv:2305.12063},
  year   = {2023}
}
R2 v1 2026-06-28T10:39:50.424Z