SoGAR:基于自监督时空注意力的社交群体活动识别
计算机视觉与模式识别
2024-11-20 v4
摘要
本文介绍了一种利用自监督 Transformer 网络进行社交群体活动识别 (SoGAR) 的新方法,该网络能有效利用无标签视频数据。为提取时空信息,我们创建了具有不同帧率的局部与全局视图。我们的自监督目标确保从同一视频的对比视图中提取的特征在时空域上保持一致。我们提出的方法高效使用基于 Transformer 的编码器来缓解群体活动识别的弱监督设定。通过利用 Transformer 模型的优势,我们的方法能对沿时空维度的长程关系进行建模。我们提出的 SoGAR 方法在 JRDB-PAR、NBA 和 Volleyball 三个群体活动识别基准上取得了 SOTA 结果,在 F1-score、MCA 和 MPCA 指标上超越了当前数值。
引用
@article{arxiv.2305.06310,
title = {SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition},
author = {Naga VS Raviteja Chappa and Pha Nguyen and Alexander H Nelson and Han-Seok Seo and Xin Li and Page Daniel Dobbs and Khoa Luu},
journal= {arXiv preprint arXiv:2305.06310},
year = {2024}
}
备注
Under review for IEEE Access journal; 12 pages, 7 figures. arXiv admin note: text overlap with arXiv:2303.12149