Spatio-Temporal Video Representation Learning for AI Based Video Playback Style Prediction
Abstract
Ever-increasing smartphone-generated video content demands intelligent techniques to edit and enhance videos on power-constrained devices. Most of the best performing algorithms for video understanding tasks like action recognition, localization, etc., rely heavily on rich spatio-temporal representations to make accurate predictions. For effective learning of the spatio-temporal representation, it is crucial to understand the underlying object motion patterns present in the video. In this paper, we propose a novel approach for understanding object motions via motion type classification. The proposed motion type classifier predicts a motion type for the video based on the trajectories of the objects present. Our classifier assigns a motion type for the given video from the following five primitive motion classes: linear, projectile, oscillatory, local and random. We demonstrate that the representations learned from the motion type classification generalizes well for the challenging downstream task of video retrieval. Further, we proposed a recommendation system for video playback style based on the motion type classifier predictions.
Cite
@article{arxiv.2110.01015,
title = {Spatio-Temporal Video Representation Learning for AI Based Video Playback Style Prediction},
author = {Rishubh Parihar and Gaurav Ramola and Ranajit Saha and Ravi Kini and Aniket Rege and Sudha Velusamy},
journal= {arXiv preprint arXiv:2110.01015},
year = {2021}
}
Comments
10 pages, 5 figures, 4 tables, ICCV Workshops 2021 - SRVU