English
Related papers

Related papers: The ActivityNet Large-Scale Activity Recognition C…

200 papers

We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker recognition technology is able to diarise and recognize…

The key prerequisite for accessing the huge potential of current machine learning techniques is the availability of large databases that capture the complex relations of interest. Previous datasets are focused on either 3D scene…

In this paper, we introduce a novel large-scale video dataset dubbed MM-SEAL for multi-person multi-grained spatio-temporal action localization among human daily life. We are the first to propose a new benchmark for multi-person…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Shimin Chen , Wei Li , Chen Chen , Jianyang Gu , Jiaming Chu , Xunqiang Tao , Yandong Guo

Human Activity Recognition (HAR) enables context-aware user experiences where mobile apps can alter content and interactions depending on user activities. Hence, smartphones have become valuable for HAR as they allow large, and diversified…

Human-Computer Interaction · Computer Science 2023-01-18 Emma Bouton--Bessac , Lakmal Meegahapola , Daniel Gatica-Perez

Visual-based human action recognition can be found in various application fields, e.g., surveillance systems, sports analytics, medical assistive technologies, or human-robot interaction frameworks, and it concerns the identification and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Antonios Gasteratos , Stavros N. Moutsis , Konstantinos A. Tsintotas , Yiannis Aloimonos

In this paper we consider the problem of classifying fine-grained, multi-step activities (e.g., cooking different recipes, making disparate home improvements, creating various forms of arts and crafts) from long videos spanning up to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Xudong Lin , Fabio Petroni , Gedas Bertasius , Marcus Rohrbach , Shih-Fu Chang , Lorenzo Torresani

Complex human activity recognition (CHAR) remains a pivotal challenge within ubiquitous computing, especially in the context of smart environments. Existing studies typically require meticulous labeling of both atomic and complex…

Artificial Intelligence · Computer Science 2024-08-07 Yuan Sun , Navid Salami Pargoo , Taqiya Ehsan , Zhao Zhang , Jorge Ortiz

Creating and labelling datasets of videos for use in training Human Activity Recognition models is an arduous task. In this paper, we approach this by using 3D rendering tools to generate a synthetic dataset of videos, and show that a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Ollie Matthews , Koki Ryu , Tarun Srivastava

Existing computer vision technologies in artwork recognition focus mainly on instance retrieval or coarse-grained attribute classification. In this work, we present a novel dataset for fine-grained artwork attribute recognition. The images…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Chenyang Zhang , Christine Kaeser-Chen , Grace Vesom , Jennie Choi , Maria Kessler , Serge Belongie

Activity recognition has become a popular research branch in the field of pervasive computing in recent years. A large number of experiments can be obtained that activity sensor-based data's characteristic in activity recognition is…

Computer Vision and Pattern Recognition · Computer Science 2018-05-21 Li Xue , Si Xiandong , Nie Lanshun , Li Jiazhen , Ding Renjie , Zhan Dechen , Chu Dianhui

We present an authentic learning task designed for computing students, centred on the creation of data-art visualisations from chosen datasets for a public exhibition. This exhibition was showcased in the cinema foyer for two weeks in June,…

Human-Computer Interaction · Computer Science 2024-08-15 Jonathan C. Roberts

A new large-scale video dataset for human action recognition, called STAIR Actions is introduced. STAIR Actions contains 100 categories of action labels representing fine-grained everyday home actions so that it can be applied to research…

Computer Vision and Pattern Recognition · Computer Science 2018-04-17 Yuya Yoshikawa , Jiaqing Lin , Akikazu Takeuchi

Partially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content. The pursuit here is of both effective and efficient solutions to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Peipei Song , Long Zhang , Long Lan , Weidong Chen , Dan Guo , Xun Yang , Meng Wang

We collected a new dataset that includes approximately eight hours of audiovisual recordings of a group of students and their self-evaluation scores for classroom engagement. The dataset and data analysis scripts are available on our…

Human-Computer Interaction · Computer Science 2023-04-19 Alpay Sabuncuoglu , T. Metin Sezgin

A key requirement for leveraging supervised deep learning methods is the availability of large, labeled datasets. Unfortunately, in the context of RGB-D scene understanding, very little data is available -- current datasets cover a small…

Computer Vision and Pattern Recognition · Computer Science 2017-04-12 Angela Dai , Angel X. Chang , Manolis Savva , Maciej Halber , Thomas Funkhouser , Matthias Nießner

In recent years, deep neural network approaches have naturally extended to the video domain, in their simplest case by aggregating per-frame classifications as a baseline for action recognition. A majority of the work in this area extends…

Computer Vision and Pattern Recognition · Computer Science 2018-01-24 Daniel Castro , Steven Hickson , Patsorn Sangkloy , Bhavishya Mittal , Sean Dai , James Hays , Irfan Essa

This report presents SceneNet and KnowledgeNet, our approaches developed for the HD-EPIC VQA Challenge 2025. SceneNet leverages scene graphs generated with a multi-modal large language model (MLLM) to capture fine-grained object…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Agnese Taluzzi , Davide Gesualdi , Riccardo Santambrogio , Chiara Plizzari , Francesca Palermo , Simone Mentasti , Matteo Matteucci

We report on CMU Informedia Lab's system used in Google's YouTube 8 Million Video Understanding Challenge. In this multi-label video classification task, our pipeline achieved 84.675% and 84.662% GAP on our evaluation split and the official…

Computer Vision and Pattern Recognition · Computer Science 2017-07-26 Po-Yao Huang , Ye Yuan , Zhenzhong Lan , Lu Jiang , Alexander G. Hauptmann

We address temporal localization of events in large-scale video data, in the context of the Youtube-8M Segments dataset. This emerging field within video recognition can enable applications to identify the precise time a specified event…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Mikel Bober-Irizar , Miha Skalic , David Austin

With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embodied agents that can operate in diverse environments given…

Robotics · Computer Science 2024-11-07 Haochen Zhang , Nader Zantout , Pujith Kachana , Zongyuan Wu , Ji Zhang , Wenshan Wang