English

Learning Priors of Human Motion With Vision Transformers

Computer Vision and Pattern Recognition 2025-01-31 v1 Robotics

Abstract

A clear understanding of where humans move in a scenario, their usual paths and speeds, and where they stop, is very important for different applications, such as mobility studies in urban areas or robot navigation tasks within human-populated environments. We propose in this article, a neural architecture based on Vision Transformers (ViTs) to provide this information. This solution can arguably capture spatial correlations more effectively than Convolutional Neural Networks (CNNs). In the paper, we describe the methodology and proposed neural architecture and show the experiments' results with a standard dataset. We show that the proposed ViT architecture improves the metrics compared to a method based on a CNN.

Keywords

Cite

@article{arxiv.2501.18543,
  title  = {Learning Priors of Human Motion With Vision Transformers},
  author = {Placido Falqueto and Alberto Sanfeliu and Luigi Palopoli and Daniele Fontanelli},
  journal= {arXiv preprint arXiv:2501.18543},
  year   = {2025}
}

Comments

2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 2024

R2 v1 2026-06-28T21:26:05.544Z