English
Related papers

Related papers: Referring Atomic Video Action Recognition

200 papers

Human action recognition (HAR) in videos is a fundamental research topic in computer vision. It consists mainly in understanding actions performed by humans based on a sequence of visual observations. In recent years, HAR have witnessed…

Computer Vision and Pattern Recognition · Computer Science 2020-10-30 Soufiane Lamghari , Guillaume-Alexandre Bilodeau , Nicolas Saunier

Video anomaly understanding (VAU) aims to provide detailed interpretation and semantic comprehension of anomalous events within videos, addressing limitations of traditional methods that focus solely on detecting and localizing anomalies.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Ying Cheng , Yu-Ho Lin , Min-Hung Chen , Fu-En Yang , Shang-Hong Lai

This paper proposes Attribute Attention Network (AANet), a new architecture that integrates person attributes and attribute attention maps into a classification framework to solve the person re-identification (re-ID) problem. Many person…

Computer Vision and Pattern Recognition · Computer Science 2019-12-20 Chiat-Pin Tay , Sharmili Roy , Kim-Hui Yap

Skeletal Action recognition from an egocentric view is important for applications such as interfaces in AR/VR glasses and human-robot interaction, where the device has limited resources. Most of the existing skeletal action recognition…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Junan Lin , Zhichao Sun , Enjie Cao , Taein Kwon , Mahdi Rad , Marc Pollefeys

Monitoring the movement and actions of humans in video in real-time is an important task. We present a deep learning based algorithm for human action recognition for both RGB and thermal cameras. It is able to detect and track humans and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Hannes Fassold , Karlheinz Gutjahr , Anna Weber , Roland Perko

Dramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still an enormous performance gap between artificial vision systems and…

Computer Vision and Pattern Recognition · Computer Science 2019-03-08 Chi Zhang , Feng Gao , Baoxiong Jia , Yixin Zhu , Song-Chun Zhu

Existing Vision-Language-Action (VLA) models can be broadly categorized into diffusion-based and auto-regressive (AR) approaches: diffusion models capture continuous action distributions but rely on computationally heavy iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Huaihai Lyu , Chaofan Chen , Senwei Xie , Pengwei Wang , Xiansheng Chen , Shanghang Zhang , Changsheng Xu

Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the…

Factorization methods for recommender systems tend to represent users as a single latent vector. However, user behavior and interests may change in the context of the recommendations that are presented to the user. For example, in the case…

Information Retrieval · Computer Science 2020-04-21 Oren Barkan , Avi Caciularu , Ori Katz , Noam Koenigstein

Video action detection (VAD) aims to detect actors and classify their actions in a video. We figure that VAD suffers more from classification rather than localization of actors. Hence, we analyze how prevailing methods form features for…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Jinsung Lee , Taeoh Kim , Inwoong Lee , Minho Shim , Dongyoon Wee , Minsu Cho , Suha Kwak

Repetitive action counting quantifies the frequency of specific actions performed by individuals. However, existing action-counting datasets have limited action diversity, potentially hampering model performance on unseen actions. To…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Jiada Lu , WeiWei Zhou , Xiang Qian , Dongze Lian , Yanyu Xu , Weifeng Wang , Lina Cao , Shenghua Gao

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by…

Machine Learning · Computer Science 2023-04-06 Alexandros Haliassos , Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Maja Pantic

Every minute, hundreds of hours of video are uploaded to social media sites and the Internet from around the world. This material creates a visual record of the experiences of a significant percentage of humanity and can help illuminate how…

Computer Vision and Pattern Recognition · Computer Science 2019-07-08 Junwei Liang , Jay D. Aronson , Alexander Hauptmann

Human activity recognition (HAR) research has increased in recent years due to its applications in mobile health monitoring, activity recognition, and patient rehabilitation. The typical approach is training a HAR classifier offline with…

Signal Processing · Electrical Eng. & Systems 2021-02-24 Sizhe An , Ganapati Bhat , Suat Gumussoy , Umit Ogras

The goal of video-based person re-identification is to match two input videos, so that the distance of the two videos is small if two videos contain the same person. A common approach for person re-identification is to first extract image…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Tanzila Rahman , Mrigank Rochan , Yang Wang

Referring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. The main concept involves learning an accurate alignment of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Baoli Sun , Xinzhu Ma , Ning Wang , Zhihui Wang , Zhiyong Wang

In this paper, we propose a deep convolutional recurrent neural network that predicts action sequences for task and motion planning (TAMP) from an initial scene image. Typical TAMP problems are formalized by combining reasoning on a…

Machine Learning · Computer Science 2020-06-11 Danny Driess , Jung-Su Ha , Marc Toussaint

Human Activity Recognition (HAR) underpins applications in healthcare, rehabilitation, fitness tracking, and smart environments, yet existing deep learning approaches demand dataset-specific training, large labeled corpora, and significant…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Nirhoshan Sivaroopan , Hansi Karunarathna , Chamara Madarasingha , Anura Jayasumana , Kanchana Thilakarathna

Facial Expression Recognition (FER) plays a crucial role in human affective analysis and has been widely applied in computer vision tasks such as human-computer interaction and psychological assessment. The 8th Affective Behavior Analysis…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 JunGyu Lee , Kunyoung Lee , Haesol Park , Ig-Jae Kim , Gi Pyo Nam

Action recognition has become a rapidly developing research field within the last decade. But with the increasing demand for large scale data, the need of hand annotated data for the training becomes more and more impractical. One way to…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Hilde Kuehne , Alexander Richard , Juergen Gall