English

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

Robotics 2026-04-24 v2 Artificial Intelligence Computation and Language Human-Computer Interaction Machine Learning

Abstract

Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different adaptations benefit from different interaction modalities. We present an interactive framework that enables robot skill adaptation through three complementary modalities: kinesthetic touch for precise spatial corrections, natural language for high-level semantic modifications, and a graphical web interface for visualizing geometric relations and trajectories, inspecting and adjusting parameters, and editing via-points by drag-and-drop. The framework integrates five components: energy-based human-intention detection, a tool-based LLM architecture (where the LLM selects and parameterizes predefined functions rather than generating code) for safe natural language adaptation, Kernelized Movement Primitives (KMPs) for motion encoding, probabilistic Virtual Fixtures for guided demonstration recording, and ergodic control for surface finishing. We demonstrate that this tool-based LLM architecture generalizes skill adaptation from KMPs to ergodic control, enabling voice-commanded surface finishing. Validation on a 7-DoF torque-controlled robot at the Automatica 2025 trade fair demonstrates the practical applicability of our approach in industrial settings.

Keywords

Cite

@article{arxiv.2604.20468,
  title  = {MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation},
  author = {Markus Knauer and Edoardo Fiorini and Maximilian Mühlbauer and Stefan Schneyer and Promwat Angsuratanawech and Florian Samuel Lay and Timo Bachmann and Samuel Bustamante and Korbinian Nottensteiner and Freek Stulp and Alin Albu-Schäffer and João Silvério and Thomas Eiband},
  journal= {arXiv preprint arXiv:2604.20468},
  year   = {2026}
}

Comments

15 pages, 13 figures, 3 tables

R2 v1 2026-07-01T12:30:15.584Z