English

Learning reusable concepts across different egocentric video understanding tasks

Computer Vision and Pattern Recognition 2025-06-02 v1

Abstract

Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen soon, everything all at once. To endow autonomous systems with such holistic perception, learning how to correlate concepts, abstract knowledge across diverse tasks, and leverage tasks synergies when learning novel skills is essential. In this paper, we introduce Hier-EgoPack, a unified framework able to create a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed.

Keywords

Cite

@article{arxiv.2505.24690,
  title  = {Learning reusable concepts across different egocentric video understanding tasks},
  author = {Simone Alberto Peirone and Francesca Pistilli and Antonio Alliegro and Tatiana Tommasi and Giuseppe Averta},
  journal= {arXiv preprint arXiv:2505.24690},
  year   = {2025}
}

Comments

Extended abstract derived from arXiv:2502.02487. Presented at the Second Joint Egocentric Vision (EgoVis) Workshop (CVPR 2025)

R2 v1 2026-07-01T02:50:50.242Z