English

Pose-Aware Multi-Level Motion Parsing for Action Quality Assessment

Computer Vision and Pattern Recognition 2025-11-11 v1

Abstract

Human pose serves as a cornerstone of action quality assessment (AQA), where subtle spatial-temporal variations in pose often distinguish excellence from mediocrity. In high-level competitions, these nuanced differences become decisive factors in scoring. In this paper, we propose a novel multi-level motion parsing framework for AQA based on enhanced spatial-temporal pose features. On the first level, the Action-Unit Parser is designed with the help of pose extraction to achieve precise action segmentation and comprehensive local-global pose representations. On the second level, Motion Parser is used by spatial-temporal feature learning to capture pose changes and appearance details for each action-unit. Meanwhile, some special conditions other than body-related will impact action scoring, like water splash in diving. In this work, we design an additional Condition Parser to offer users more flexibility in their choices. Finally, Weight-Adjust Scoring Module is introduced to better accommodate the diverse requirements of various action types and the multi-scale nature of action-units. Extensive evaluations on large-scale diving sports datasets demonstrate that our multi-level motion parsing framework achieves state-of-the-art performance in both action segmentation and action scoring tasks.

Cite

@article{arxiv.2511.05611,
  title  = {Pose-Aware Multi-Level Motion Parsing for Action Quality Assessment},
  author = {Shuaikang Zhu and Yang Yang and Chen Sun},
  journal= {arXiv preprint arXiv:2511.05611},
  year   = {2025}
}
R2 v1 2026-07-01T07:26:55.297Z