English
Related papers

Related papers: MEVA: A Large-Scale Multiview, Multimodal Video Da…

200 papers

The main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which are suffered from illumination, fast motion, privacy-preserving, and large energy consumption. Meanwhile, the biologically inspired…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Xiao Wang , Zongzhen Wu , Bo Jiang , Zhimin Bao , Lin Zhu , Guoqi Li , Yaowei Wang , Yonghong Tian

The 3rd annual installment of the ActivityNet Large- Scale Activity Recognition Challenge, held as a full-day workshop in CVPR 2018, focused on the recognition of daily life, high-level, goal-oriented activities from user-generated videos…

Computer Vision and Pattern Recognition · Computer Science 2018-08-24 Bernard Ghanem , Juan Carlos Niebles , Cees Snoek , Fabian Caba Heilbron , Humam Alwassel , Victor Escorcia , Ranjay Krishna , Shyamal Buch , Cuong Duc Dao

Despite significant progress in the development of human action detection datasets and algorithms, no current dataset is representative of real-world aerial view scenarios. We present Okutama-Action, a new video dataset for aerial view…

Computer Vision and Pattern Recognition · Computer Science 2017-06-16 Mohammadamin Barekatain , Miquel Martí , Hsueh-Fu Shih , Samuel Murray , Kotaro Nakayama , Yutaka Matsuo , Helmut Prendinger

The Codec Avatars Lab at Meta introduces Embody 3D, a multimodal dataset of 500 individual hours of 3D motion data from 439 participants collected in a multi-camera collection stage, amounting to over 54 million frames of tracked 3D motion.…

Analyzing student actions is an important and challenging task in educational research. Existing efforts have been hampered by the lack of accessible datasets to capture the nuanced action dynamics in classrooms. In this paper, we present a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Zhuolin Tan , Chenqiang Gao , Anyong Qin , Ruixin Chen , Tiecheng Song , Feng Yang , Deyu Meng

Facial expression analysis based on machine learning requires large number of well-annotated data to reflect different changes in facial motion. Publicly available datasets truly help to accelerate research in this area by providing a…

Machine Learning · Computer Science 2023-01-31 Yanfu Yan , Ke Lu , Jian Xue , Pengcheng Gao , Jiayi Lyu

Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods. However, they waste an excessive amount of computations and memory in predicting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Sunil Hwang , Jaehong Yoon , Youngwan Lee , Sung Ju Hwang

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization between speech,…

Artificial Intelligence · Computer Science 2024-10-23 Saif Punjwani , Larry Heck

Video Annotation is a crucial process in computer science and social science alike. Many video annotation tools (VATs) offer a wide range of features for making annotation possible. We conducted an extensive survey of over 59 VATs and…

Human-Computer Interaction · Computer Science 2023-01-10 Snehesh Shrestha , William Sentosatio , Huiashu Peng , Cornelia Fermuller , Yiannis Aloimonos

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

Most existing autonomous-driving datasets (e.g., KITTI, nuScenes, and the Waymo Perception Dataset), collected by human-driving mode or unidentified driving mode, can only serve as early training for the perception and prediction of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Xiangyu Li , Chen Wang , Yumao Liu , Dengbo He , Jiahao Zhang , Ke Ma

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Time dataset (MIT). For…

Machine Learning · Computer Science 2023-08-21 Michael Joannou , Pia Rotshtein , Uta Noppeney

Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtain. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Naila Murray , Antonio Manuel López

Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments. This limits the utility of machine learning (ML) models learned from them.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Ali Zia , Renuka Sharma , Reza Arablouei , Greg Bishop-Hurley , Jody McNally , Neil Bagnall , Vivien Rolland , Brano Kusy , Lars Petersson , Aaron Ingham

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Ranjay Krishna , Kenji Hata , Frederic Ren , Li Fei-Fei , Juan Carlos Niebles

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

Surveillance videos are an essential component of daily life with various critical applications, particularly in public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Tongtong Yuan , Xuange Zhang , Kun Liu , Bo Liu , Chen Chen , Jian Jin , Zhenzhen Jiao

We introduce the AVECL-UMons dataset for audio-visual event classification and localization in the context of office environments. The audio-visual dataset is composed of 11 event classes recorded at several realistic positions in two…

Information Retrieval · Computer Science 2020-11-03 Mathilde Brousmiche , Stéphane Dupont , Jean Rouat

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as…

Computation and Language · Computer Science 2024-06-05 Zhe Chen , Heyang Liu , Wenyi Yu , Guangzhi Sun , Hongcheng Liu , Ji Wu , Chao Zhang , Yu Wang , Yanfeng Wang

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Tiantian Geng , Teng Wang , Jinming Duan , Runmin Cong , Feng Zheng