ImageInThat:通过操控图像向机器人传递用户指令
人机交互
2025-03-21 v1 机器人学
摘要
基础模型正在显著提升机器人在日常任务中的自主执行能力,如餐饮准备等,但由于模型性能、难以捕捉用户偏好以及需要用户主导性,机器人仍需依赖人类指令。现有指令方式多样:自然语言表达直接但可能抽象或模糊,而面向终端用户的编程支持长期任务,但界面难以捕捉用户意图。本研究提出以图像直接操控作为替代范式,引入具体实现方案 ImageInThat,允许用户在时间轴式界面中对图像进行直接操控,以生成机器人指令。通过用户研究,我们展示了 ImageInThat 在厨房操作任务中的效果,并与基于文本的自然语言指令方法进行比较。结果表明,参与者在 ImageInThat 下完成任务速度更快,且更倾向于使用该方法。补充材料包括代码可在:https://image-in-that.github.io/ 查阅。
引用
@article{arxiv.2503.15500,
title = {ImageInThat: Manipulating Images to Convey User Instructions to Robots},
author = {Karthik Mahadevan and Blaine Lewis and Jiannan Li and Bilge Mutlu and Anthony Tang and Tovi Grossman},
journal= {arXiv preprint arXiv:2503.15500},
year = {2025}
}
备注
In Proceedings of the ACM/IEEE International Conference on Human-Robot Interaction (HRI), 2025