English
Related papers

Related papers: The AVA-Kinetics Localized Human Actions Video Dat…

200 papers

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

Human-Computer Interaction · Computer Science 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

Current researches of action recognition mainly focus on single-view and multi-view recognition, which can hardly satisfies the requirements of human-robot interaction (HRI) applications to recognize actions from arbitrary views. The lack…

Computer Vision and Pattern Recognition · Computer Science 2019-04-25 Yanli Ji , Feixiang Xu , Yang Yang , Fumin Shen , Heng Tao Shen , Wei-Shi Zheng

In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Lingru Zhou , Yiqi Gao , Manqing Zhang , Peng Wu , Peng Wang , Yanning Zhang

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

Video anomaly retrieval aims to localize anomalous events in videos using natural language queries to facilitate public safety. However, existing datasets suffer from severe limitations: (1) data scarcity due to the long-tail nature of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Shuyu Yang , Yilun Wang , Yaxiong Wang , Li Zhu , Zhedong Zheng

We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We present an extensive…

Computer Vision and Pattern Recognition · Computer Science 2022-02-22 Oana Ignat , Santiago Castro , Yuhang Zhou , Jiajun Bao , Dandan Shan , Rada Mihalcea

We present a new dataset with annotated eye movements. The dataset consists of over 800,000 gaze points recorded during a car ride in the real world and in the simulator. In total, the eye movements of 19 subjects were annotated. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-01-13 Wolfgang Fuhl , Enkelejda Kasneci

Most existing traffic video datasets including Waymo are structured, focusing predominantly on Western traffic, which hinders global applicability. Specifically, most Asian scenarios are far more complex, involving numerous objects with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Xijun Wang , Pedro Sandoval-Segura , Chengyuan Zhang , Junyun Huang , Tianrui Guan , Ruiqi Xian , Fuxiao Liu , Rohan Chandra , Boqing Gong , Dinesh Manocha

This paper introduces a new large consent-driven dataset aimed at assisting in the evaluation of algorithmic bias and robustness of computer vision and audio speech models in regards to 11 attributes that are self-provided or labeled by…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Bilal Porgali , Vítor Albiero , Jordan Ryda , Cristian Canton Ferrer , Caner Hazirbas

Repetitive action counting quantifies the frequency of specific actions performed by individuals. However, existing action-counting datasets have limited action diversity, potentially hampering model performance on unseen actions. To…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Jiada Lu , WeiWei Zhou , Xiang Qian , Dongze Lian , Yanyu Xu , Weifeng Wang , Lina Cao , Shenghua Gao

In recent years, automatic video caption generation has attracted considerable attention. This paper focuses on the generation of Japanese captions for describing human actions. While most currently available video caption datasets have…

Computation and Language · Computer Science 2020-03-11 Yutaro Shigeto , Yuya Yoshikawa , Jiaqing Lin , Akikazu Takeuchi

In this paper, we introduce a new hierarchical model for human action recognition using body joint locations. Our model can categorize complex actions in videos, and perform spatio-temporal annotations of the atomic actions that compose the…

Computer Vision and Pattern Recognition · Computer Science 2016-06-17 Ivan Lillo , Juan Carlos Niebles , Alvaro Soto

Action recognition models have achieved impressive results by incorporating scene-level annotations, such as objects, their relations, 3D structure, and more. However, obtaining annotations of scene structure for videos requires a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Roei Herzig , Ofir Abramovich , Elad Ben-Avraham , Assaf Arbelle , Leonid Karlinsky , Ariel Shamir , Trevor Darrell , Amir Globerson

This dissertation presents a methodology for recording speed climbing training sessions with multiple cameras and annotating the videos with relevant data, including body position, hand and foot placement, and timing. The annotated data is…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Yufei Xie , Shaoman Li , Penghui Lin

Video Annotation is a crucial process in computer science and social science alike. Many video annotation tools (VATs) offer a wide range of features for making annotation possible. We conducted an extensive survey of over 59 VATs and…

Human-Computer Interaction · Computer Science 2023-01-10 Snehesh Shrestha , William Sentosatio , Huiashu Peng , Cornelia Fermuller , Yiannis Aloimonos

We introduce Latent Action Pretraining for general Action models (LAPA), an unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require…

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Rui Qian , Tao Wang , Ning Xu , Hongkai Xiong , Guo-Jun Qi , Nicu Sebe

This work presents the Industrial Hand Action Dataset V1, an industrial assembly dataset consisting of 12 classes with 459,180 images in the basic version and 2,295,900 images after spatial augmentation. Compared to other freely available…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Fabian Sturm , Elke Hergenroether , Julian Reinhardt , Petar Smilevski Vojnovikj , Melanie Siegel

In the world of action recognition research, one primary focus has been on how to construct and train networks to model the spatial-temporal volume of an input video. These methods typically uniformly sample a segment of an input clip…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Xinyu Li , Chunhui Liu , Bing Shuai , Yi Zhu , Hao Chen , Joseph Tighe

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as…

Computation and Language · Computer Science 2024-06-05 Zhe Chen , Heyang Liu , Wenyi Yu , Guangzhi Sun , Hongcheng Liu , Ji Wu , Chao Zhang , Yu Wang , Yanfeng Wang
‹ Prev 1 3 4 5 6 7 10 Next ›