English
Related papers

Related papers: A Short Note on the Kinetics-700 Human Action Data…

200 papers

Detecting 3D objects keypoints is of great interest to the areas of both graphics and computer vision. There have been several 2D and 3D keypoint datasets aiming to address this problem in a data-driven way. These datasets, however, either…

Computer Vision and Pattern Recognition · Computer Science 2020-08-10 Yang You , Yujing Lou , Chengkun Li , Zhoujun Cheng , Liangwei Li , Lizhuang Ma , Weiming Wang , Cewu Lu

This paper presents a framework to automate the labelling process for gestures in musical performance videos with a 3D Convolutional Neural Network (CNN). While this idea was proposed in a previous study, this paper introduces several…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Foteini Simistira Liwicki , Richa Upadhyay , Prakash Chandra Chhipa , Killian Murphy , Federico Visi , Stefan Östersjö , Marcus Liwicki

We address the problem of data augmentation for video action recognition. Standard augmentation strategies in video are hand-designed and sample the space of possible augmented data points either at random, without knowing which augmented…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Shreyank N Gowda , Marcus Rohrbach , Frank Keller , Laura Sevilla-Lara

Technologies play an increasingly important role in sports and become a real competitive advantage for the athletes who benefit from it. Among them, the use of motion capture is developing in various sports to optimize sporting gestures.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-09 Fiche Guénolé , Sevestre Vincent , Gonzalez-Barral Camila , Leglaive Simon , Séguier Renaud

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Time dataset (MIT). For…

Machine Learning · Computer Science 2023-08-21 Michael Joannou , Pia Rotshtein , Uta Noppeney

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jiahao Wang , Yufeng Yuan , Rujie Zheng , Youtian Lin , Jian Gao , Lin-Zhuo Chen , Yajie Bao , Yi Zhang , Chang Zeng , Yanxi Zhou , Xiao-Xiao Long , Hao Zhu , Zhaoxiang Zhang , Xun Cao , Yao Yao

We collected a new dataset that includes approximately eight hours of audiovisual recordings of a group of students and their self-evaluation scores for classroom engagement. The dataset and data analysis scripts are available on our…

Human-Computer Interaction · Computer Science 2023-04-19 Alpay Sabuncuoglu , T. Metin Sezgin

We propose a novel method for 3D point cloud action recognition. Understanding human actions in RGB videos has been widely studied in recent years, however, its 3D point cloud counterpart remains under-explored. This is mostly due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Yizhak Ben-Shabat , Oren Shrout , Stephen Gould

We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Although we have witnessed great progress for video action…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Junwei Liang , Enwei Zhang , Jun Zhang , Chunhua Shen

Current fully-supervised video datasets consist of only a few hundred thousand videos and fewer than a thousand domain-specific labels. This hinders the progress towards advanced video architectures. This paper presents an in-depth study of…

Computer Vision and Pattern Recognition · Computer Science 2019-05-03 Deepti Ghadiyaram , Matt Feiszli , Du Tran , Xueting Yan , Heng Wang , Dhruv Mahajan

We present a bundle-adjustment-based algorithm for recovering accurate 3D human pose and meshes from monocular videos. Unlike previous algorithms which operate on single frames, we show that reconstructing a person over an entire sequence…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Anurag Arnab , Carl Doersch , Andrew Zisserman

Synthesis of long-term human motion skeleton sequences is essential to aid human-centric video generation with potential applications in Augmented Reality, 3D character animations, pedestrian trajectory prediction, etc. Long-term human…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Neeraj Battan , Yudhik Agrawal , Veeravalli Saisooryarao , Aman Goel , Avinash Sharma

The current biodiversity loss crisis makes animal monitoring a relevant field of study. In light of this, data collected through monitoring can provide essential insights, and information for decision-making aimed at preserving global…

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Yansong Tang , Jinpeng Liu , Aoyang Liu , Bin Yang , Wenxun Dai , Yongming Rao , Jiwen Lu , Jie Zhou , Xiu Li

The increasing variety and quantity of tagged multimedia content on a variety of online platforms offer a unique opportunity to advance the field of human action recognition. In this study, we utilize 283,582 unique, unlabeled TikTok video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Yang Qian , Yinan Sun , Ali Kargarandehkordi , Parnian Azizian , Onur Cezmi Mutlu , Saimourya Surabhi , Pingyi Chen , Zain Jabbar , Dennis Paul Wall , Peter Washington

Classifying the behavior of humans or animals from videos is important in biomedical fields for understanding brain function and response to stimuli. Action recognition, classifying activities performed by one or more subjects in a trimmed…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Michael Perez , Corey Toler-Franklin

In this paper, we introduce RoleMotion, a large-scale human motion dataset that encompasses a wealth of role-playing and functional motion data tailored to fit various specific scenes. Existing text datasets are mainly constructed…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Junran Peng , Yiheng Huang , Silei Shen , Zeji Wei , Jingwei Yang , Baojie Wang , Yonghao He , Chuanchen Luo , Man Zhang , Xucheng Yin , Wei Sui

MVImgNet is a large-scale dataset that contains multi-view images of ~220k real-world objects in 238 classes. As a counterpart of ImageNet, it introduces 3D visual signals via multi-view shooting, making a soft bridge between 2D and 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Xiaoguang Han , Yushuang Wu , Luyue Shi , Haolin Liu , Hongjie Liao , Lingteng Qiu , Weihao Yuan , Xiaodong Gu , Zilong Dong , Shuguang Cui

In this paper, we study the value of using synthetically produced videos as training data for neural networks used for action categorization. Motivated by the fact that texture and background of a video play little to no significant roles…

Computer Vision and Pattern Recognition · Computer Science 2020-01-31 Mohamad Ballout , Mohammad Tuqan , Daniel Asmar , Elie Shammas , George Sakr

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi