English
Related papers

Related papers: An Evaluation of Action Recognition Models on EPIC…

200 papers

In this work, we introduce our solution to the EPIC-KITCHENS-100 2022 Action Detection challenge. One-stage Action Detection Transformer (OADT) is proposed to model the temporal connection of video segments. With the help of OADT, both the…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lijun Li , Li'an Zhuo , Bang Zhang

Despite the growing discriminative capabilities of modern deep learning methods for recognition tasks, the inner workings of the state-of-art models still remain mostly black-boxes. In this paper, we propose a systematic interpretation of…

Computer Vision and Pattern Recognition · Computer Science 2017-11-27 Jingxuan Hou , Tae Soo Kim , Austin Reiter

Human activity recognition using deep learning techniques has become increasing popular because of its high effectivity with recognizing complex tasks, as well as being relatively low in costs compared to more traditional machine learning…

Computer Vision and Pattern Recognition · Computer Science 2022-04-29 Wei Zhong Tee , Rushit Dave , Naeem Seliya , Mounika Vanamala

Action recognition is an important and challenging problem in video analysis. Although the past decade has witnessed progress in action recognition with the development of deep learning, such process has been slow in competitive sports…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Shenlan Liu , Xiang Liu , Gao Huang , Lin Feng , Lianyu Hu , Dong Jiang , Aibin Zhang , Yang Liu , Hong Qiao

We introduce Activity Graph Transformer, an end-to-end learnable model for temporal action localization, that receives a video as input and directly predicts a set of action instances that appear in the video. Detecting and localizing…

Computer Vision and Pattern Recognition · Computer Science 2021-01-29 Megha Nawhal , Greg Mori

This paper is a brief report to our submission to the VIPriors Action Recognition Challenge. Action recognition has attracted many researchers attention for its full application, but it is still challenging. In this paper, we study previous…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Zhipeng Luo , Dawei Xu , Zhiguang Zhang

Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Pulkit Kumar , Shuaiyi Huang , Matthew Walmer , Sai Saketh Rambhatla , Abhinav Shrivastava

Despite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the action recognition task…

Computer Vision and Pattern Recognition · Computer Science 2020-07-03 Ziyang Song , Ziyi Yin , Zejian Yuan , Chong Zhang , Wanchao Chi , Yonggen Ling , Shenghao Zhang

Current video/action understanding systems have demonstrated impressive performance on large recognition tasks. However, they might be limiting themselves to learning to recognize spatiotemporal patterns, rather than attempting to…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Paritosh Parmar , Brendan Morris

This paper focuses on task recognition and action segmentation in weakly-labeled instructional videos, where only the ordered sequence of video-level actions is available during training. We propose a two-stream framework, which exploits…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Reza Ghoddoosian , Saif Sayed , Vassilis Athitsos

In this paper, we study current and upcoming frontiers across the landscape of skeleton-based human action recognition. To study skeleton-action recognition in the wild, we introduce Skeletics-152, a curated and 3-D pose-annotated subset of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Pranay Gupta , Anirudh Thatipelli , Aditya Aggarwal , Shubh Maheshwari , Neel Trivedi , Sourav Das , Ravi Kiran Sarvadevabhatla

We introduce LEAP (illustrated in Figure 1), a novel method for generating video-grounded action programs through use of a Large Language Model (LLM). These action programs represent the motoric, perceptual, and structural aspects of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Eadom Dessalene , Michael Maynord , Cornelia Fermüller , Yiannis Aloimonos

Deep learning has achieved remarkable success in modeling sequential data, including event sequences, temporal point processes, and irregular time series. Recently, transformers have largely replaced recurrent networks in these tasks.…

Machine Learning · Computer Science 2025-08-05 Ivan Karpukhin , Andrey Savchenko

This paper studies introducing viewpoint invariant feature representations in existing action recognition architecture. Despite significant progress in action recognition, efficiently handling geometric variations in large-scale datasets…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Jinhui Ye , Junwei Liang

Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Raiyaan Abdullah , Jared Claypoole , Michael Cogswell , Ajay Divakaran , Yogesh Rawat

Foundation models (FMs) are large neural networks trained on broad datasets, excelling in downstream tasks with minimal fine-tuning. Human activity recognition in video has advanced with FMs, driven by competition among different…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Thinesh Thiyakesan Ponbagavathi , Kunyu Peng , Alina Roitberg

Inspired by recent advances in neural machine translation, that jointly align and translate using encoder-decoder networks equipped with attention, we propose an attentionbased LSTM model for human activity recognition. Our model jointly…

Computer Vision and Pattern Recognition · Computer Science 2017-09-01 Atousa Torabi , Leonid Sigal

Egocentric action anticipation aims to predict the future actions the camera wearer will perform from the observation of the past. While predictions about the future should be available before the predicted events take place, most…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Antonino Furnari , Giovanni Maria Farinella

Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Andy Bonnetto , Haozhe Qi , Franklin Leong , Matea Tashkovska , Mahdi Rad , Solaiman Shokur , Friedhelm Hummel , Silvestro Micera , Marc Pollefeys , Alexander Mathis

Research on depth-based human activity analysis achieved outstanding performance and demonstrated the effectiveness of 3D representation for action recognition. The existing depth-based and RGB+D-based action recognition benchmarks have a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Jun Liu , Amir Shahroudy , Mauricio Perez , Gang Wang , Ling-Yu Duan , Alex C. Kot
‹ Prev 1 3 4 5 6 7 10 Next ›