English
Related papers

Related papers: The EPIC-KITCHENS Dataset: Collection, Challenges …

200 papers

In this paper, we approach an open problem of artwork identification and propose a new dataset dubbed Open Museum Identification Challenge (Open MIC). It contains photos of exhibits captured in 10 distinct exhibition spaces of several…

Computer Vision and Pattern Recognition · Computer Science 2018-02-06 Piotr Koniusz , Yusuf Tas , Hongguang Zhang , Mehrtash Harandi , Fatih Porikli , Rui Zhang

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Existing activity tracker datasets for human activity recognition are typically obtained by having participants perform predefined activities in an enclosed environment under supervision. This results in small datasets with a limited number…

Human-Computer Interaction · Computer Science 2024-03-01 Shing Chan , Hang Yuan , Catherine Tong , Aidan Acquah , Abram Schonfeldt , Jonathan Gershuny , Aiden Doherty

Our lives can be seen as a complex weaving of activities; we switch from one activity to another, to maximise our achievements or in reaction to demands placed upon us. Observing a video of unscripted daily activities, we parse the video…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Will Price , Carl Vondrick , Dima Damen

Videos can evoke a range of affective responses in viewers. The ability to predict evoked affect from a video, before viewers watch the video, can help in content creation and video recommendation. We introduce the Evoked Expressions from…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Jennifer J. Sun , Ting Liu , Alan S. Cowen , Florian Schroff , Hartwig Adam , Gautam Prasad

Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Cheng-Yen Yang , Hsiang-Wei Huang , Zhongyu Jiang , Hao Wang , Farron Wallace , Jenq-Neng Hwang

Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physical processes remains difficult, as models must combine…

We introduce a new large-scale data set of video URLs with densely-sampled object bounding box annotations called YouTube-BoundingBoxes (YT-BB). The data set consists of approximately 380,000 video segments about 19s long, automatically…

Computer Vision and Pattern Recognition · Computer Science 2017-03-28 Esteban Real , Jonathon Shlens , Stefano Mazzocchi , Xin Pan , Vincent Vanhoucke

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Karttikeya Mangalam , Raiymbek Akshulakov , Jitendra Malik

Advances in deep generative modeling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactive datasets that…

Machine Learning · Computer Science 2026-02-19 Yingchen He , Christian D. Weilbach , Martyna E. Wojciechowska , Yuxuan Zhang , Frank Wood

Falls are significant and often fatal for vulnerable populations such as the elderly. Previous works have addressed the detection of falls by relying on data capture by a single sensor, images or accelerometers. In this work, we rely on…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Xueyi Wang

Automatic detection of intake gestures is a key element of automatic dietary monitoring. Several types of sensors, including inertial measurement units (IMU) and video cameras, have been used for this purpose. The common machine learning…

Human-Computer Interaction · Computer Science 2020-10-01 Philipp V. Rouast , Hamid Heydarian , Marc T. P. Adam , Megan E. Rollo

We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions…

We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible research in web agents. It contains 31,725 trajectories and 318k steps, featuring a core…

Artificial Intelligence · Computer Science 2026-04-15 Sicheng Fan , Rui Wan , Yifei Leng , Gaoning Liang , Li Ling , Yanyi Shang , Dehan Kong

Emergency Medical Services (EMS) are critical to patient survival in emergencies, but first responders often face intense cognitive demands in high-stakes situations. AI cognitive assistants, acting as virtual partners, have the potential…

In Brain-Computer Interface (BCI) research, the detailed study of blinks is crucial. They can be considered as noise, affecting the efficiency and accuracy of decoding users' cognitive states and intentions, or as potential features,…

Neurons and Cognition · Quantitative Biology 2025-06-10 E. Guttmann-Flury , X. Sheng , X. Zhu

We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered scenarios. PACE provides a large-scale real-world benchmark…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Yang You , Kai Xiong , Zhening Yang , Zhengxiang Huang , Junwei Zhou , Ruoxi Shi , Zhou Fang , Adam W. Harley , Leonidas Guibas , Cewu Lu

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-body action…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Yuan-Ming Li , Wei-Jin Huang , An-Lan Wang , Ling-An Zeng , Jing-Ke Meng , Wei-Shi Zheng

Camera-based passive dietary intake monitoring is able to continuously capture the eating episodes of a subject, recording rich visual information, such as the type and volume of food being consumed, as well as the eating behaviours of the…

Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Ranjay Krishna , Kenji Hata , Frederic Ren , Li Fei-Fei , Juan Carlos Niebles