English

OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization

Computer Vision and Pattern Recognition 2026-07-01 v1

Abstract

Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We introduce Point-Supervised Online TAL (POTAL), which localizes actions in streaming videos using only one temporal point per instance. To solve POTAL, we propose OnPoint, an offline-to-online multi-level distillation framework that transfers knowledge from a point-supervised offline teacher to an online student via (i) pseudo-segment instance distillation, (ii) class-activation sequence distillation, and (iii) anticipatory window-level distillation. We further improve robustness by incorporating the original point labels into student training and by refining anchor decoding with actionness-guided attention calibration. Experiments on five datasets show OnPoint consistently outperforms strong baselines, establishing a solid foundation for POTAL.

Keywords

Cite

@article{arxiv.2607.00289,
  title  = {OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization},
  author = {Sakib Reza and Gauri Jagatap and Mohsen Moghaddam and Octavia Camps and Andrea Fanelli},
  journal= {arXiv preprint arXiv:2607.00289},
  year   = {2026}
}

Comments

Accepted at ECCV 2026