English
Related papers

Related papers: PKU-MMD: A Large Scale Benchmark for Continuous Mu…

200 papers

We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions…

A new large-scale video dataset for human action recognition, called STAIR Actions is introduced. STAIR Actions contains 100 categories of action labels representing fine-grained everyday home actions so that it can be applied to research…

Computer Vision and Pattern Recognition · Computer Science 2018-04-17 Yuya Yoshikawa , Jiaqing Lin , Akikazu Takeuchi

In this paper, we present an approach for identification of actions within depth action videos. First, we process the video to get motion history images (MHIs) and static history images (SHIs) corresponding to an action video based on the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Mohammad Farhad Bulbul , Saiful Islam , Hazrat Ali

Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Fan Bai , Yuxin Du , Tiejun Huang , Max Q. -H. Meng , Bo Zhao

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots' dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University…

Robotics · Computer Science 2024-07-02 Yifan Tang , Cong Tai , Fangxing Chen , Wanting Zhang , Tao Zhang , Xueping Liu , Yongjin Liu , Long Zeng

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Riko Suzuki , Hitomi Yanaka , Koji Mineshima , Daisuke Bekki

Periodic human activities with implicit workflows are common in manufacturing, sports, and daily life. While short-term periodic activities -- characterized by simple structures and high-contrast patterns -- have been widely studied,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Fan Yang , Quanting Xie , Atsunori Moteki , Shoichi Masui , Shan Jiang , Kanji Uchino , Yonatan Bisk , Graham Neubig

This work focuses on per-video unsupervised action segmentation, which is of interest to applications where storing large datasets is either not possible, or nor permitted. We propose to segment videos by learning in deep kernel space, to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Silvia L. Pintea , Jouke Dijkstra

We present DogMo, a large-scale multi-view RGB-D video dataset capturing diverse canine movements for the task of motion recovery from images. DogMo comprises 1.2k motion sequences collected from 10 unique dogs, offering rich variation in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Zan Wang , Siyu Chen , Luya Mo , Xinfeng Gao , Yuxin Shen , Lebin Ding , Wei Liang

Accurate analysis of combat sports using computer vision has gained traction in recent years, yet the development of robust datasets remains a major bottleneck due to the dynamic, unstructured nature of actions and variations in recording…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Rahul Kumar , Vipul Baghel , Sudhanshu Singh , Bikash Kumar Badatya , Shivam Yadav , Babji Srinivasan , Ravi Hegde

The fine-grained action analysis of the existing action datasets is challenged by insufficient action categories, low fine granularities, limited modalities, and tasks. In this paper, we propose a Multi-modality and Multi-task dataset of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Sheng-Lan Liu , Yu-Ning Ding , Gang Yan , Si-Fan Zhang , Jin-Rong Zhang , Wen-Yue Chen , Xue-Hai Xu

Real-time 3D human action recognition has broad industrial applications, such as surveillance, human-computer interaction, and healthcare monitoring. By relying on complex spatio-temporal local encoding, most existing point cloud sequence…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Xing Li , Qian Huang , Zhijian Wang , Zhenjie Hou , Tianjin Yang , Zhuang Miao

Moments capture a huge part of our lives. Accurate recognition of these moments is challenging due to the diverse and complex interpretation of the moments. Action recognition refers to the act of classifying the desired action/activity…

Computer Vision and Pattern Recognition · Computer Science 2018-09-14 Ankit Shah , Harini Kesavamoorthy , Poorva Rane , Pramati Kalwad , Alexander Hauptmann , Florian Metze

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

Computer Vision and Pattern Recognition · Computer Science 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Human activity recognition is one of the most important tasks in computer vision and has proved useful in different fields such as healthcare, sports training and security. There are a number of approaches that have been explored to solve…

Computer Vision and Pattern Recognition · Computer Science 2023-05-01 Sheryl Mathew , Annapoorani Subramanian , Pooja , Balamurugan MS , Manoj Kumar Rajagopal

This study uses multisensory data (i.e., color and depth) to recognize human actions in the context of multimodal human-robot interaction. Here we employed the iCub robot to observe the predefined actions of the human partners by using four…

Robotics · Computer Science 2022-12-20 Kas Kniesmeijer , Murat Kirtay

Recently, deep learning approach has achieved promising results in various fields of computer vision. In this paper, a new framework called Hierarchical Depth Motion Maps (HDMM) + 3 Channel Deep Convolutional Neural Networks (3ConvNets) is…

Computer Vision and Pattern Recognition · Computer Science 2015-01-21 Pichao Wang , Wanqing Li , Zhimin Gao , Jing Zhang , Chang Tang , Philip Ogunbona

In this paper, we propose a novel technique for measuring behavioral engagement through students' actions recognition. The proposed approach recognizes student actions then predicts the student behavioral engagement level. For student…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Ahmed Abdelkawy , Aly Farag , Islam Alkabbany , Asem Ali , Chris Foreman , Thomas Tretter , Nicholas Hindy

A great number of computer vision publications have focused on distinguishing between human action recognition and classification rather than the intensity of actions performed. Indexing the intensity which determines the performance of…

Artificial Intelligence · Computer Science 2020-03-27 Nihar Bendre , Nima Ebadi , John J Prevost , Paul Rad

Human action Recognition for unknown views is a challenging task. We propose a view-invariant deep human action recognition framework, which is a novel integration of two important action cues: motion and shape temporal dynamics (STD). The…

Computer Vision and Pattern Recognition · Computer Science 2020-01-22 Chhavi Dhiman , Dinesh Kumar Vishwakarma