English
Related papers

Related papers: ActionArt: Advancing Multimodal Large Models for F…

200 papers

Due to burdensome data requirements, learning from demonstration often falls short of its promise to allow users to quickly and naturally program robots. Demonstrations are inherently ambiguous and incomplete, making correct generalization…

Machine Learning · Computer Science 2019-04-29 Wonjoon Goo , Scott Niekum

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Zihui Xue , Kristen Grauman

Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates on fully observed videos, action anticipation must handle…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Seulgi Kim , Ghazal Kaviani , Mohit Prabhushankar , Ghassan AlRegib

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

Videos are more informative than images because they capture the dynamics of the scene. By representing motion in videos, we can capture dynamic activities. In this work, we introduce GPT-4 generated motion descriptions that capture…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Chinmaya Devaraj , Cornelia Fermuller , Yiannis Aloimonos

Fine-grained action recognition is attracting increasing attention due to the emerging demand of specific action understanding in real-world applications, whereas the data of rare fine-grained categories is very limited. Therefore, we…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Jiahao Wang , Yunhong Wang , Sheng Liu , Annan Li

Due to the rapid temporal and fine-grained nature of complex human assembly atomic actions, traditional action segmentation approaches requiring the spatial (and often temporal) down sampling of video frames often loose vital fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Matthew Kent Myers , Nick Wright , Stephen McGough , Nicholas Martin

Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Baoli Sun , Yihan Wang , Xinzhu Ma , Zhihui Wang , Kun Lu , Zhiyong Wang

The advent of Multimodal LLMs has significantly enhanced image OCR recognition capabilities, making GUI automation a viable reality for increasing efficiency in digital tasks. One fundamental aspect of developing a GUI automation system is…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Qinchen Wu , Difei Gao , Kevin Qinghong Lin , Zhuoyu Wu , Xiangwu Guo , Peiran Li , Weichen Zhang , Hengxu Wang , Mike Zheng Shou

Existing action quality assessment (AQA) methods mainly learn deep representations at the video level for scoring diverse actions. Due to the lack of a fine-grained understanding of actions in videos, they harshly suffer from low…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Jinglin Xu , Sibo Yin , Guohao Zhao , Zishuo Wang , Yuxin Peng

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Katsuyuki Nakamura , Hiroki Ohashi , Mitsuhiro Okada

Humans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Xuan Ju , Ailing Zeng , Jianan Wang , Qiang Xu , Lei Zhang

With advances in data-driven machine learning research, a wide variety of prediction models have been proposed to capture spatio-temporal features for the analysis of video streams. Recognising actions and detecting action transitions…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Harshala Gammulle , David Ahmedt-Aristizabal , Simon Denman , Lachlan Tychsen-Smith , Lars Petersson , Clinton Fookes

Despite the rapid progress, existing works on action understanding focus strictly on one type of action agent, which we call actor---a human adult, ignoring the diversity of actions performed by other actors. To overcome this narrow…

Computer Vision and Pattern Recognition · Computer Science 2017-05-01 Chenliang Xu , Caiming Xiong , Jason J. Corso

Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved remarkable success in representation learning. Due to image-level visual-language…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Mengxiao Tian , Xinxiao Wu , Shuo Yang

We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun action labels, by…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Ivan Rodin , Antonino Furnari , Kyle Min , Subarna Tripathi , Giovanni Maria Farinella

Our world is full of varied actions and moves across specialized domains that we, as humans, strive to identify and understand. Within any single domain, actions can often appear quite similar, making it challenging for deep models to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Mohammadreza Salehi , Jae Sung Park , Tanush Yadav , Aditya Kusupati , Ranjay Krishna , Yejin Choi , Hannaneh Hajishirzi , Ali Farhadi

Visual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limitations of existing datasets: (1) insufficient annotation granularity, which impedes…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Minghao Zou , Qingtian Zeng , Yongping Miao , Shangkun Liu , Zilong Wang , Hantao Liu , Wei Zhou

Despite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the action recognition task…

Computer Vision and Pattern Recognition · Computer Science 2020-07-03 Ziyang Song , Ziyi Yin , Zejian Yuan , Chong Zhang , Wanchao Chi , Yonggen Ling , Shenghao Zhang