English
Related papers

Related papers: MEVA: A Large-Scale Multiview, Multimodal Video Da…

200 papers

We present Aria Everyday Activities (AEA) Dataset, an egocentric multimodal open dataset recorded using Project Aria glasses. AEA contains 143 daily activity sequences recorded by multiple wearers in five geographically diverse indoor…

The Out the Window (OTW) dataset is a crowdsourced activity dataset containing 5,668 instances of 17 activities from the NIST Activities in Extended Video (ActEV) challenge. These videos are crowdsourced from workers on the Amazon…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Gregory Castanon , Nathan Shnidman , Tim Anderson , Jeffrey Byrne

MEx: Multi-modal Exercises Dataset is a multi-sensor, multi-modal dataset, implemented to benchmark Human Activity Recognition(HAR) and Multi-modal Fusion algorithms. Collection of this dataset was inspired by the need for recognising and…

Computer Vision and Pattern Recognition · Computer Science 2019-08-27 Anjana Wijekoon , Nirmalie Wiratunga , Kay Cooper

We make available to the community a new dataset to support action-recognition research. This dataset is different from prior datasets in several key ways. It is significantly larger. It contains streaming video with long segments…

Computer Vision and Pattern Recognition · Computer Science 2015-11-19 Daniel Paul Barrett , Ran Xu , Haonan Yu , Jeffrey Mark Siskind

Unmanned aerial vehicles (UAVs) are widely applied for purposes of inspection, search, and rescue operations by the virtue of low-cost, large-coverage, real-time, and high-resolution data acquisition capacities. Massive volumes of aerial…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Pu Jin , Lichao Mou , Gui-Song Xia , Xiao Xiang Zhu

Understanding human behavior from complementary egocentric (ego) and exocentric (exo) points of view enables the development of systems that can support workers in industrial environments and enhance their safety. However, progress in this…

Deep learning for human action recognition in videos is making significant progress, but is slowed down by its dependency on expensive manual labeling of large video collections. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2017-07-20 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Antonio Manuel López Peña

Body-worn first-person vision (FPV) camera enables to extract a rich source of information on the environment from the subject's viewpoint. However, the research progress in wearable camera-based egocentric office activity understanding is…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Girmaw Abebe Tadesse , Oliver Bent , Komminist Weldemariam , Md. Abrar Istiak , Taufiq Hasan , Andrea Cavallaro

Recently, event-based vision sensors have gained attention for autonomous driving applications, as conventional RGB cameras face limitations in handling challenging dynamic conditions. However, the availability of real-world and synthetic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Manideep Reddy Aliminati , Bharatesh Chakravarthi , Aayush Atul Verma , Arpitsinh Vaghela , Hua Wei , Xuesong Zhou , Yezhou Yang

Modern vehicles equip dashcams that primarily collect visual evidence for traffic accidents. However, most of the video data collected by dashcams that is not related to traffic accidents is discarded without any use. In this paper, we…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-12-02 Seyul Lee , Jayden King , Young Choon Lee , Hyuck Han , Sooyong Kang

Comprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB, LiDAR, or IMU data. However, solely…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Ming Yan , Yan Zhang , Shuqiang Cai , Shuqi Fan , Xincheng Lin , Yudi Dai , Siqi Shen , Chenglu Wen , Lan Xu , Yuexin Ma , Cheng Wang

Humans are arguably one of the most important subjects in video streams, many real-world applications such as video summarization or video editing workflows often require the automatic search and retrieval of a person of interest. Despite…

Computer Vision and Pattern Recognition · Computer Science 2021-06-04 Juan Leon Alcazar , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem , Fabian Caba Heilbron

We present a dataset with models of 14 articulated objects commonly found in human environments and with RGB-D video sequences and wrenches recorded of human interactions with them. The 358 interaction sequences total 67 minutes of human…

Robotics · Computer Science 2018-06-19 Roberto Martín-Martín , Clemens Eppner , Oliver Brock

Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the variability and complexity of human activities. Despite…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Reuben Tan , Matthias De Lange , Michael Iuzzolino , Bryan A. Plummer , Kate Saenko , Karl Ridgeway , Lorenzo Torresani

We live in a world filled with never-ending streams of multimodal information. As a more natural recording of the real scenario, long form audio-visual videos are expected as an important bridge for better exploring and understanding the…

Multimedia · Computer Science 2023-06-19 Wenxuan Hou , Guangyao Li , Yapeng Tian , Di Hu

Video Anomaly Detection (VAD) finds widespread applications in security surveillance, traffic monitoring, industrial monitoring, and healthcare. Despite extensive research efforts, there remains a lack of concise reviews that provide…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Liyun Zhu , Lei Wang , Arjun Raj , Tom Gedeon , Chen Chen

We present NERVE (Neuromorphic Vision and Radar Ensemble), a multi-sensor dataset comprising 257 minutes of synchronized recordings from five sensors: two Dynamic Vision Sensors (DVS), an RGB-D camera, and two Radar units (24GHz and 77GHz).…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Omar Mansour , Pietro Martinello , Ethan Milon , YingFu Xu , Manolis Sifalakis , Guangzhi Tang , Amirreza Yousefzadeh

Event detection (ED), which means identifying event trigger words and classifying event types, is the first and most fundamental step for extracting event knowledge from plain text. Most existing datasets exhibit the following issues that…

Computation and Language · Computer Science 2020-10-09 Xiaozhi Wang , Ziqi Wang , Xu Han , Wangyi Jiang , Rong Han , Zhiyuan Liu , Juanzi Li , Peng Li , Yankai Lin , Jie Zhou

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jiahao Wang , Yufeng Yuan , Rujie Zheng , Youtian Lin , Jian Gao , Lin-Zhuo Chen , Yajie Bao , Yi Zhang , Chang Zeng , Yanxi Zhou , Xiao-Xiao Long , Hao Zhu , Zhaoxiang Zhang , Xun Cao , Yao Yao

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples…

Computer Vision and Pattern Recognition · Computer Science 2016-07-28 Gunnar A. Sigurdsson , Gül Varol , Xiaolong Wang , Ali Farhadi , Ivan Laptev , Abhinav Gupta