English

Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport

Computer Vision and Pattern Recognition 2025-11-13 v2

Abstract

We study self-supervised procedure learning, which discovers key steps and their order from a set of unlabeled videos. Previous methods typically learn frame-to-frame correspondences between videos before determining key steps and their order. However, their performance often suffers from order variations, background/redundant frames, and repeated actions. To overcome these challenges, we propose a self-supervised framework, which utilizes a fused Gromov-Wasserstein optimal transport with a structural prior for frame-to-frame mapping. However, optimizing only for the above temporal alignment may lead to degenerate solutions, where all frames are mapped to a small cluster in the embedding space and thus every video is assigned to just one key step. To address that issue, we integrate a contrastive regularization, which maps different frames to various points, avoiding trivial solutions. Finally, extensive experiments on egocentric and third-person benchmarks demonstrate our superior performance over prior works, including OPEL which relies on a classical Kantorovich optimal transport with an optimality prior.

Keywords

Cite

@article{arxiv.2507.15540,
  title  = {Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport},
  author = {Syed Ahmed Mahmood and Ali Shah Ali and Umer Ahmed and Fawad Javed Fateh and M. Zeeshan Zia and Quoc-Huy Tran},
  journal= {arXiv preprint arXiv:2507.15540},
  year   = {2025}
}

Comments

Accepted to WACV 2026

R2 v1 2026-07-01T04:11:10.993Z