中文

面向导盲机器人的 Space-Aware 指令调优:数据集与基准

机器人学 2025-02-13 v2 计算机视觉与模式识别

摘要

导盲机器人为增强视障者的 mobility 和 safety 提供了有前景的解决方案, addresses 传统导盲犬在 perceptual intelligence 和 communication 方面的局限性。随着 Vision-Language Models (VLMs) 的出现,机器人现在能够 generate natural language 描述其周围环境, aid 在更安全的 decision-making 中。然而,现有的 VLMs 常常 struggle to accurately interpret and convey spatial relationships,这对于在复杂环境中如街道十字路口的 navigation 至关重要。我们引入 Space-Aware 指令调优 (SAIT) 数据集和 Space-Aware 基准 (SA-Bench) 来 addresses 当前 VLMs 在理解 physical 环境方面的局限性。我们的 automated 数据生成管道 focus on virtual path to destination 在 3D 空间中的路径以及周围环境, enhance 环境理解能力, enable VLMs 提供更 accurate 指导给视障者。我们还提出了 evaluate VLM 在 deliver walking guidance 方面 effectiveness 的 protocol。比较实验表明,我们的 space-aware instruction-tuned 模型优于 state-of-the-art 算法。我们已在 https://github.com/byungokhan/Space-awareVLM 完全 open-source 了 SAIT 数据集和 SA-Bench 以及相关 code。

关键词

引用

@article{arxiv.2502.07183,
  title  = {Space-Aware Instruction Tuning: Dataset and Benchmark for Guide Dog Robots Assisting the Visually Impaired},
  author = {ByungOk Han and Woo-han Yun and Beom-Su Seo and Jaehong Kim},
  journal= {arXiv preprint arXiv:2502.07183},
  year   = {2025}
}

备注

ICRA 2025