LIT:面向大语言模型驱动的意图跟踪用于主动人类-机器人协作——基于机器人主厨的应用
机器人学
2024-06-21 v1 计算机视觉与模式识别
摘要
大语言模型(LLM)和视觉语言模型(VLM)使机器人能够将自然语言提示锚定到控制动作中,以实现开放世界中的任务。然而,对于长期跨越的协作任务而言,这种表述会导致在任务每个步骤都需要频繁调用以发起或澄清机器人动作。我们提出语言驱动意图跟踪(LIT),利用 LLM 和 VLM 模型用户的长期行为,以预测下一个用户意图以引导机器人进行主动协作。我们演示了基于 LIT 的协作机器人与用户在烹饪协作任务中的顺畅协调。
引用
@article{arxiv.2406.13787,
title = {LIT: Large Language Model Driven Intention Tracking for Proactive Human-Robot Collaboration -- A Robot Sous-Chef Application},
author = {Zhe Huang and John Pohovey and Ananya Yammanuru and Katherine Driggs-Campbell},
journal= {arXiv preprint arXiv:2406.13787},
year = {2024}
}
备注
Spotlight Presentation at the 3rd Workshop on Computer Vision in the Wild at CVPR 2024. Also accepted by the 5th Annual Embodied AI Workshop at CVPR 2024