中文
相关论文

相关论文: MEVA: A Large-Scale Multiview, Multimodal Video Da…

200 篇论文

The main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which are suffered from illumination, fast motion, privacy-preserving, and large energy consumption. Meanwhile, the biologically inspired…

计算机视觉与模式识别 · 计算机科学 2022-11-18 Xiao Wang , Zongzhen Wu , Bo Jiang , Zhimin Bao , Lin Zhu , Guoqi Li , Yaowei Wang , Yonghong Tian

The 3rd annual installment of the ActivityNet Large- Scale Activity Recognition Challenge, held as a full-day workshop in CVPR 2018, focused on the recognition of daily life, high-level, goal-oriented activities from user-generated videos…

Despite significant progress in the development of human action detection datasets and algorithms, no current dataset is representative of real-world aerial view scenarios. We present Okutama-Action, a new video dataset for aerial view…

计算机视觉与模式识别 · 计算机科学 2017-06-16 Mohammadamin Barekatain , Miquel Martí , Hsueh-Fu Shih , Samuel Murray , Kotaro Nakayama , Yutaka Matsuo , Helmut Prendinger

The Codec Avatars Lab at Meta introduces Embody 3D, a multimodal dataset of 500 individual hours of 3D motion data from 439 participants collected in a multi-camera collection stage, amounting to over 54 million frames of tracked 3D motion.…

Analyzing student actions is an important and challenging task in educational research. Existing efforts have been hampered by the lack of accessible datasets to capture the nuanced action dynamics in classrooms. In this paper, we present a…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Zhuolin Tan , Chenqiang Gao , Anyong Qin , Ruixin Chen , Tiecheng Song , Feng Yang , Deyu Meng

Facial expression analysis based on machine learning requires large number of well-annotated data to reflect different changes in facial motion. Publicly available datasets truly help to accelerate research in this area by providing a…

机器学习 · 计算机科学 2023-01-31 Yanfu Yan , Ke Lu , Jian Xue , Pengcheng Gao , Jiayi Lyu

Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods. However, they waste an excessive amount of computations and memory in predicting…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Sunil Hwang , Jaehong Yoon , Youngwan Lee , Sung Ju Hwang

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization between speech,…

人工智能 · 计算机科学 2024-10-23 Saif Punjwani , Larry Heck

Video Annotation is a crucial process in computer science and social science alike. Many video annotation tools (VATs) offer a wide range of features for making annotation possible. We conducted an extensive survey of over 59 VATs and…

人机交互 · 计算机科学 2023-01-10 Snehesh Shrestha , William Sentosatio , Huiashu Peng , Cornelia Fermuller , Yiannis Aloimonos

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

Most existing autonomous-driving datasets (e.g., KITTI, nuScenes, and the Waymo Perception Dataset), collected by human-driving mode or unidentified driving mode, can only serve as early training for the perception and prediction of…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Xiangyu Li , Chen Wang , Yumao Liu , Dengbo He , Jiahao Zhang , Ke Ma

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Time dataset (MIT). For…

机器学习 · 计算机科学 2023-08-21 Michael Joannou , Pia Rotshtein , Uta Noppeney

Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtain. In this work, we investigate the generation of synthetic…

计算机视觉与模式识别 · 计算机科学 2019-10-16 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Naila Murray , Antonio Manuel López

Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments. This limits the utility of machine learning (ML) models learned from them.…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Ali Zia , Renuka Sharma , Reza Arablouei , Greg Bishop-Hurley , Jody McNally , Neil Bagnall , Vivien Rolland , Brano Kusy , Lars Petersson , Aaron Ingham

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both…

计算机视觉与模式识别 · 计算机科学 2017-05-03 Ranjay Krishna , Kenji Hata , Frederic Ren , Li Fei-Fei , Juan Carlos Niebles

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

计算机视觉与模式识别 · 计算机科学 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

Surveillance videos are an essential component of daily life with various critical applications, particularly in public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events.…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Tongtong Yuan , Xuange Zhang , Kun Liu , Bo Liu , Chen Chen , Jian Jin , Zhenzhen Jiao

We introduce the AVECL-UMons dataset for audio-visual event classification and localization in the context of office environments. The audio-visual dataset is composed of 11 event classes recorded at several realistic positions in two…

信息检索 · 计算机科学 2020-11-03 Mathilde Brousmiche , Stéphane Dupont , Jean Rouat

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as…

计算与语言 · 计算机科学 2024-06-05 Zhe Chen , Heyang Liu , Wenyi Yu , Guangzhi Sun , Hongcheng Liu , Ji Wu , Chao Zhang , Yu Wang , Yanfeng Wang

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Tiantian Geng , Teng Wang , Jinming Duan , Runmin Cong , Feng Zheng