English
Related papers

Related papers: Referring Atomic Video Action Recognition

200 papers

Despite excellent progress has been made, the performance on action recognition still heavily relies on specific datasets, which are difficult to extend new action classes due to labor-intensive labeling. Moreover, the high diversity in…

Computer Vision and Pattern Recognition · Computer Science 2020-11-18 Xiaoyuan Ni , Sizhe Song , Yu-Wing Tai , Chi-Keung Tang

Localizing persons and recognizing their actions from videos is a challenging task towards high-level video understanding. Recent advances have been achieved by modeling direct pairwise relations between entities. In this paper, we take one…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Junting Pan , Siyu Chen , Mike Zheng Shou , Yu Liu , Jing Shao , Hongsheng Li

Visual dialog is a challenging vision-language task, which requires the agent to answer multi-round questions about an image. It typically needs to address two major problems: (1) How to answer visually-grounded questions, which is the core…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Yulei Niu , Hanwang Zhang , Manli Zhang , Jianhong Zhang , Zhiwu Lu , Ji-Rong Wen

Text-based person search aims to retrieve specific individuals across camera networks using natural language descriptions. However, current benchmarks often exhibit biases towards common actions like walking or standing, neglecting the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Shuyu Yang , Yaxiong Wang , Li Zhu , Zhedong Zheng

Action recognition is a key problem in computer vision that labels videos with a set of predefined actions. Capturing both, semantic content and motion, along the video frames is key to achieve high accuracy performance on this task. Most…

Computer Vision and Pattern Recognition · Computer Science 2019-10-23 Xia Huang , Hossein Mousavi , Gemma Roig

Part-level Action Parsing aims at part state parsing for boosting action recognition in videos. Despite of dramatic progresses in the area of video classification research, a severe problem faced by the community is that the detailed…

Computer Vision and Pattern Recognition · Computer Science 2021-11-08 Xuanhan Wang , Xiaojia Chen , Lianli Gao , Lechao Chen , Jingkuan Song

This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these…

Computer Vision and Pattern Recognition · Computer Science 2020-05-21 Ang Li , Meghana Thotakuri , David A. Ross , João Carreira , Alexander Vostrikov , Andrew Zisserman

This work focuses on the apparent emotional reaction recognition (AERR) from the video-only input, conducted in a self-supervised fashion. The network is first pre-trained on different self-supervised pretext tasks and later fine-tuned on…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Marija Jegorova , Stavros Petridis , Maja Pantic

A large amount of recent research has focused on tasks that combine language and vision, resulting in a proliferation of datasets and methods. One such task is action recognition, whose applications include image annotation, scene under-…

Computation and Language · Computer Science 2017-04-25 Spandana Gella , Frank Keller

We propose two well-motivated ranking-based methods to enhance the performance of current state-of-the-art human activity recognition systems. First, as an improvement over the classic power normalization method, we propose a parameter-free…

Computer Vision and Pattern Recognition · Computer Science 2015-12-14 Zhenzhong Lan , Shoou-I Yu , Alexander G. Hauptmann

The Segment Anything Model (SAM) has gained significant attention for its impressive performance in image segmentation. However, it lacks proficiency in referring video object segmentation (RVOS) due to the need for precise user-interactive…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Yonglin Li , Jing Zhang , Xiao Teng , Long Lan , Xinwang Liu

Action recognition is a critical task for social robots to meaningfully engage with their environment. 3D human skeleton-based action recognition is an attractive research area in recent years. Although, the existing approaches are good at…

Computer Vision and Pattern Recognition · Computer Science 2021-01-01 Hui Feng , Shanshan Wang , Shuzhi Sam Ge

Research in human action recognition has accelerated significantly since the introduction of powerful machine learning tools such as Convolutional Neural Networks (CNNs). However, effective and efficient methods for incorporation of…

Computer Vision and Pattern Recognition · Computer Science 2018-03-21 Jinliang Zang , Le Wang , Ziyi Liu , Qilin Zhang , Zhenxing Niu , Gang Hua , Nanning Zheng

Procedure planning requires a model to predict a sequence of actions that transform a start visual observation into a goal in instructional videos. While most existing methods rely primarily on visual observations as input, they often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Lei Shi , Victor Aregbede , Andreas Persson , Martin Längkvist , Amy Loutfi , Stephanie Lowry

Emotion recognition can provide crucial information about the user in many applications when building human-computer interaction (HCI) systems. Most of current researches on visual emotion recognition are focusing on exploring facial…

Computer Vision and Pattern Recognition · Computer Science 2018-05-31 Man-Chin Sun , Shih-Huan Hsu , Min-Chun Yang , Jen-Hsien Chien

Action parsing in videos with complex scenes is an interesting but challenging task in computer vision. In this paper, we propose a generic 3D convolutional neural network in a multi-task learning manner for effective Deep Action Parsing…

Computer Vision and Pattern Recognition · Computer Science 2016-02-11 Li Liu , Yi Zhou , Ling Shao

We present a method to generate video-action pairs that follow text instructions, starting from an initial image observation and the robot's joint states. Our approach automatically provides action labels for video diffusion models,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Liudi Yang , Yang Bai , George Eskandar , Fengyi Shen , Mohammad Altillawi , Dong Chen , Ziyuan Liu , Abhinav Valada

Human drivers focus only on a handful of agents at any one time. On the other hand, autonomous driving systems process complex scenes with numerous agents, regardless of whether they are pedestrians on a crosswalk or vehicles parked on the…

Machine Learning · Computer Science 2025-09-25 Carlo Bosio , Greg Woelki , Noureldin Hendy , Nicholas Roy , Byungsoo Kim

Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then…

Robotics · Computer Science 2025-09-30 Chaoran Zhu , Hengyi Wang , Yik Lung Pang , Changjae Oh

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has many applications,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 James Burgess , Xiaohan Wang , Yuhui Zhang , Anita Rau , Alejandro Lozano , Lisa Dunlap , Trevor Darrell , Serena Yeung-Levy
‹ Prev 1 4 5 6 7 8 10 Next ›