English
Related papers

Related papers: The AVA-Kinetics Localized Human Actions Video Dat…

200 papers

This paper presents our solution to the AVA-Kinetics Crossover Challenge of ActivityNet workshop at CVPR 2021. Our solution utilizes multiple types of relation modeling methods for spatio-temporal action detection and adopts a training…

Computer Vision and Pattern Recognition · Computer Science 2021-06-17 Yutong Feng , Jianwen Jiang , Ziyuan Huang , Zhiwu Qing , Xiang Wang , Shiwei Zhang , Mingqian Tang , Yue Gao

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Traditional VQA methods,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Yanyun Pu , Kehan Li , Zeyi Huang , Zhijie Zhong , Kaixiang Yang

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

Computer Vision and Pattern Recognition · Computer Science 2022-05-12 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

Short-form videos (SVs) have become a vital part of our online routine for acquiring and sharing information. Their multimodal complexity poses new challenges for video analysis, highlighting the need for video emotion analysis (VEA) within…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xuecheng Wu , Dingkang Yang , Danlei Huang , Xinyi Yin , Yifan Wang , Jia Zhang , Jiayu Nie , Liangyu Fu , Yang Liu , Junxiao Xue , Hadi Amirpour , Wei Zhou

Automatic detection of natural disasters and incidents has become more important as a tool for fast response. There have been many studies to detect incidents using still images and text. However, the number of approaches that exploit…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Duygu Sesver , Alp Eren Gençoğlu , Çağrı Emre Yıldız , Zehra Günindi , Faeze Habibi , Ziya Ata Yazıcı , Hazım Kemal Ekenel

Manual spatio-temporal annotation of human action in videos is laborious, requires several annotators and contains human biases. In this paper, we present a weakly supervised approach to automatically obtain spatio-temporal annotations of…

Computer Vision and Pattern Recognition · Computer Science 2016-05-27 Waqas Sultani , Mubarak Shah

With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions during assembly. To…

Robotics · Computer Science 2023-04-18 Dustin Aganian , Benedict Stephan , Markus Eisenbach , Corinna Stretz , Horst-Michael Gross

Human movements are both an area of intense study and the basis of many applications such as character animation. For many applications, it is crucial to identify movements from videos or analyze datasets of movements. Here we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2021-07-14 Saeed Ghorbani , Kimia Mahdaviani , Anne Thaler , Konrad Kording , Douglas James Cook , Gunnar Blohm , Nikolaus F. Troje

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However, their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Cheng-You Lu , Peisen Zhou , Angela Xing , Chandradeep Pokhariya , Arnab Dey , Ishaan Shah , Rugved Mavidipalli , Dylan Hu , Andrew Comport , Kefan Chen , Srinath Sridhar

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

Computer Vision and Pattern Recognition · Computer Science 2021-08-13 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

We introduce a simple baseline for action localization on the AVA dataset. The model builds upon the Faster R-CNN bounding box detection framework, adapted to operate on pure spatiotemporal features - in our case produced exclusively by an…

Computer Vision and Pattern Recognition · Computer Science 2018-07-27 Rohit Girdhar , João Carreira , Carl Doersch , Andrew Zisserman

The objective of action quality assessment is to score sports videos. However, most existing works focus only on video dynamic information (i.e., motion information) but ignore the specific postures that an athlete is performing in a video,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Ling-An Zeng , Fa-Ting Hong , Wei-Shi Zheng , Qi-Zhi Yu , Wei Zeng , Yao-Wei Wang , Jian-Huang Lai

Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry. To enable technological breakthrough, we present HA-ViD - the first human assembly video dataset that features representative…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Hao Zheng , Regina Lee , Yuqian Lu

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot…

Lane detection plays a key role in autonomous driving. While car cameras always take streaming videos on the way, current lane detection works mainly focus on individual images (frames) by ignoring dynamics along the video. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Yujun Zhang , Lei Zhu , Wei Feng , Huazhu Fu , Mingqian Wang , Qingxia Li , Cheng Li , Song Wang

Video Object Segmentation (VOS) is crucial for several applications, from video editing to video data generation. Training a VOS model requires an abundance of manually labeled training videos. The de-facto traditional way of annotating…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Thanos Delatolas , Vicky Kalogeiton , Dim P. Papadopoulos

Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Dewen Zhang , Tahir Hussain , Wangpeng An , Hayaru Shouno

We introduce the first audio-visual dataset for traffic anomaly detection taken from real-world scenes, called MAVAD, with a diverse range of weather and illumination conditions. In addition, we propose a novel method named AVACA that…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Błażej Leporowski , Arian Bakhtiarnia , Nicole Bonnici , Adrian Muscat , Luca Zanella , Yiming Wang , Alexandros Iosifidis

We propose a method for human action recognition, one that can localize the spatiotemporal regions that `define' the actions. This is a challenging task due to the subtlety of human actions in video and the co-occurrence of contextual…

Computer Vision and Pattern Recognition · Computer Science 2019-04-12 Yang Wang , Vinh Tran , Gedas Bertasius , Lorenzo Torresani , Minh Hoai

Vision-Language-Action (VLA) models have shown remarkable progress in embodied tasks recently, but most methods process visual observations independently at each timestep. This history-agnostic design treats robot manipulation as a Markov…

Machine Learning · Computer Science 2026-04-13 Lei Xiao , Jifeng Li , Juntao Gao , Feiyang Ye , Yan Jin , Jingjing Qian , Jing Zhang , Yong Wu , Xiaoyuan Yu