中文

使用视觉 Transformer 学习人类运动先验

计算机视觉与模式识别 2025-01-31 v1 机器人学

摘要

对人类在特定情境中移动位置、常用路径与速度以及停止位置的清晰认识,对不同应用(如城市交通研究或人类密集环境中的机器人导航任务)至关重要。本文提出一种基于视觉 Transformer(ViT)的神经网络架构,以提供上述信息。该方案可谱系比基于卷积神经网络(CNN)更有效地捕捉空间关联性。本文描述了方法与 proposed 神经网络架构,并展示了基于标准数据集的实验结果。我们表明,所提出的 ViT 架构在指标上优于基于 CNN 的方法。

关键词

引用

@article{arxiv.2501.18543,
  title  = {Learning Priors of Human Motion With Vision Transformers},
  author = {Placido Falqueto and Alberto Sanfeliu and Luigi Palopoli and Daniele Fontanelli},
  journal= {arXiv preprint arXiv:2501.18543},
  year   = {2025}
}

备注

2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 2024