中文
相关论文

相关论文: The Kinetics Human Action Video Dataset

200 篇论文

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has many applications,…

计算机视觉与模式识别 · 计算机科学 2025-03-12 James Burgess , Xiaohan Wang , Yuhui Zhang , Anita Rau , Alejandro Lozano , Lisa Dunlap , Trevor Darrell , Serena Yeung-Levy

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Delong Chen , Tejaswi Kasarla , Yejin Bang , Mustafa Shukor , Willy Chung , Jade Yu , Allen Bolourchi , Theo Moutakanni , Pascale Fung

Human Action Recognition is an important task of Human Robot Interaction as cooperation between robots and humans requires that artificial agents recognise complex cues from the environment. A promising approach is using trained classifiers…

计算机视觉与模式识别 · 计算机科学 2019-08-26 Frederico Belmonte Klein , Angelo Cangelosi

Machine learning and computer vision methods have a major impact on the study of natural animal behavior, as they enable the (semi-)automatic analysis of vast amounts of video data. Mice are the standard mammalian model system in most…

In recent years, we have seen an emergence of data-driven approaches in robotics. However, most existing efforts and datasets are either in simulation or focus on a single task in isolation such as grasping, pushing or poking. In order to…

机器人学 · 计算机科学 2018-10-17 Pratyusha Sharma , Lekha Mohan , Lerrel Pinto , Abhinav Gupta

State-of-the-art temporal action detectors inefficiently search the entire video for specific actions. Despite the encouraging progress these methods achieve, it is crucial to design automated approaches that only explore parts of the video…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Humam Alwassel , Fabian Caba Heilbron , Bernard Ghanem

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

计算机视觉与模式识别 · 计算机科学 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

The advancement of computer vision and machine learning has made datasets a crucial element for further research and applications. However, the creation and development of robots with advanced recognition capabilities are hindered by the…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Zhengcheng Shen , Yi Gao , Linh Kästner , Jens Lambrecht

Due to the statistical complexity of video, the high degree of inherent stochasticity, and the sheer amount of data, generating natural video remains a challenging task. State-of-the-art video generation models often attempt to address…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Dirk Weissenborn , Oscar Täckström , Jakob Uszkoreit

Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in…

Human movement analysis is a key area of research in robotics, biomechanics, and data science. It encompasses tracking, posture estimation, and movement synthesis. While numerous methodologies have evolved over time, a systematic and…

机器人学 · 计算机科学 2023-05-11 Brenda Elizabeth Olivas-Padilla , Alina Glushkova , Sotiris Manitsaris

Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption dataset designed to advance research in human-centric…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Yi-Xing Peng , Qize Yang , Yu-Ming Tang , Shenghao Fu , Kun-Yu Lin , Xihan Wei , Wei-Shi Zheng

We aim to automatically identify human action reasons in online videos. We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them. We introduce and make publicly available the WhyAct…

计算机视觉与模式识别 · 计算机科学 2021-09-10 Oana Ignat , Santiago Castro , Hanwen Miao , Weiji Li , Rada Mihalcea

Human movements are both an area of intense study and the basis of many applications such as character animation. For many applications, it is crucial to identify movements from videos or analyze datasets of movements. Here we introduce a…

计算机视觉与模式识别 · 计算机科学 2021-07-14 Saeed Ghorbani , Kimia Mahdaviani , Anne Thaler , Konrad Kording , Douglas James Cook , Gunnar Blohm , Nikolaus F. Troje

There is a large variation in the activities that humans perform in their everyday lives. We consider modeling these composite human activities which comprises multiple basic level actions in a completely unsupervised setting. Our model…

计算机视觉与模式识别 · 计算机科学 2016-03-14 Chenxia Wu , Jiemi Zhang , Ozan Sener , Bart Selman , Silvio Savarese , Ashutosh Saxena

In this paper, we present an approach for identification of actions within depth action videos. First, we process the video to get motion history images (MHIs) and static history images (SHIs) corresponding to an action video based on the…

计算机视觉与模式识别 · 计算机科学 2019-04-02 Mohammad Farhad Bulbul , Saiful Islam , Hazrat Ali

Human action recognition has become one of the most active field of research in computer vision due to its wide range of applications, like surveillance, medical, industrial environments, smart homes, among others. Recently, deep learning…

计算机视觉与模式识别 · 计算机科学 2020-12-29 Samuel Felipe dos Santos , Jurandy Almeida

Recently, NVS in human-object interaction scenes has received increasing attention. Existing human-object interaction datasets mainly consist of static data with limited views, offering only RGB images or videos, mostly containing…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Shuai Guo , Houqiang Zhong , Qiuwen Wang , Ziyu Chen , Yijie Gao , Jiajing Yuan , Chenyu Zhang , Rong Xie , Li Song

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Yixuan Ye , Xuanyu Lu , Yuxin Jiang , Yuchao Gu , Rui Zhao , Qiwei Liang , Jiachun Pan , Fengda Zhang , Weijia Wu , Alex Jinpeng Wang

Recent works on dynamic 3D neural field reconstruction assume the input from synchronized multi-view videos whose poses are known. The input constraints are often not satisfied in real-world setups, making the approach impractical. We show…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Changwoon Choi , Jeongjun Kim , Geonho Cha , Minkwan Kim , Dongyoon Wee , Young Min Kim