English

SportSkills: Physical Skill Learning from Sports Instructional Videos

Computer Vision and Pattern Recognition 2026-03-27 v1

Abstract

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared towards physical skill learning with in-the-wild video. SportSkills has more than 360k instructional videos containing more than 630k visual demonstrations paired with instructional narrations explaining the know-how behind the actions from 55 varied sports. Through a suite of experiments, we show that SportSkills unlocks the ability to understand fine-grained differences between physical actions. Our representation achieves gains of up to 4x with the same model trained on traditional activity-centric datasets. Crucially, building on SportSkills, we introduce the first large-scale task formulation of mistake-conditioned instructional video retrieval, bridging representation learning and actionable feedback generation (e.g., "here's my execution of a skill; which video clip should I watch to improve it?"). Formal evaluations by professional coaches show our retrieval approach significantly advances the ability of video models to personalize visual instructions for a user query.

Keywords

Cite

@article{arxiv.2603.25163,
  title  = {SportSkills: Physical Skill Learning from Sports Instructional Videos},
  author = {Kumar Ashutosh and Chi Hsuan Wu and Kristen Grauman},
  journal= {arXiv preprint arXiv:2603.25163},
  year   = {2026}
}

Comments

Technical report