Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding
Abstract
Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting native event streams with dense temporal annotations is costly. We introduce Event ActivityNet, a large-scale simulated-event benchmark derived from human-annotated, untrimmed ActivityNet videos. It comprises 3,263 videos, 200 action classes, and 106.94 hours, with matched 5-bin and 9-bin event-voxel representations, temporal action annotations, and timestamped captions. The benchmark supports annotated-segment action recognition, auxiliary event-language alignment, and causal online temporal action localization. We generate event voxels directly from non-interpolated source videos in decoded frame order, retain per-video rational nominal or average frame-rate metadata for approximate time mapping, and use action-center reconstruction LPIPS as a soft diagnostic of retained reconstructable content. We establish baselines for adaptive event framing, prompt-caption alignment, and event-only, RGB-only, and RGB-event localization. Under a progressive nested-scale training protocol, recognition Top-1 accuracy increases from 52.25 to 66.42, while online temporal localization average mAP improves from 21.7 to 29.0. Moreover, staged Event ActivityNet pretraining followed by native-event fine-tuning consistently outperforms target-only and joint-from-scratch training across multiple supervision budgets. Event ActivityNet provides a scalable benchmark for long-horizon event modeling, although native-camera evaluation remains essential for deployment-oriented conclusions.
Keywords
Cite
@article{arxiv.2608.01948,
title = {Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding},
author = {Cheng-Yao Hong and Ting-Wei Lin and Yun-Chung Lai and Hua-Wei Lee and Hwann-Tzong Chen and Tyng-Luh Liu},
journal= {arXiv preprint arXiv:2608.01948},
year = {2026}
}
Comments
22 pages, 10 figures