人类-机器人交互中多模态感知、语言 grounding 与控制的消失实验:基于目标检测与抓取任务
机器人学
2026-05-05 v1 人工智能
摘要
本稿件扩展了我们之前的多模态人类-机器人交互系统,通过引入对三个最强影响端到端性能模块的受控消失研究:用于动作提取的大型语言模型、用于视觉 grounding 的感知系统以及用于运动执行的控制器。目标并非重新设计完整管道,而是要在共同实验协议下隔离每个组件的贡献,然后对最佳组合进行端到端评估。因此,我们比较了三个语言模型、五种感知配置和三个控制器,随后进行第二阶段的阶乘研究,以确定最佳候选者。 resulting analysis is intended to clarify which choices primarily affect execution time, which primarily affect success rate, and where the largest engineering gains are likely to come from in future revisions of the system.
引用
@article{arxiv.2605.00963,
title = {Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task},
author = {Zi Tian and Guanting Shen},
journal= {arXiv preprint arXiv:2605.00963},
year = {2026}
}
备注
10 pages