English
Related papers

Related papers: COIN: A Large-scale Dataset for Comprehensive Inst…

200 papers

This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing…

Computer Vision and Pattern Recognition · Computer Science 2021-09-20 Dima Damen , Hazel Doughty , Giovanni Maria Farinella , Antonino Furnari , Evangelos Kazakos , Jian Ma , Davide Moltisanti , Jonathan Munro , Toby Perrett , Will Price , Michael Wray

Instruction tuning represents a prevalent strategy employed by Multimodal Large Language Models (MLLMs) to align with human instructions and adapt to new tasks. Nevertheless, MLLMs encounter the challenge of adapting to users' evolving…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Cheng Chen , Junchen Zhu , Xu Luo , Hengtao Shen , Lianli Gao , Jingkuan Song

In recent years, deep neural network approaches have naturally extended to the video domain, in their simplest case by aggregating per-frame classifications as a baseline for action recognition. A majority of the work in this area extends…

Computer Vision and Pattern Recognition · Computer Science 2018-01-24 Daniel Castro , Steven Hickson , Patsorn Sangkloy , Bhavishya Mittal , Sean Dai , James Hays , Irfan Essa

Despite the number of currently available datasets on video question answering, there still remains a need for a dataset involving multi-step and non-factoid answers. Moreover, relying on video transcripts remains an under-explored topic.…

Computation and Language · Computer Science 2020-06-02 Anthony Colas , Seokhwan Kim , Franck Dernoncourt , Siddhesh Gupte , Daisy Zhe Wang , Doo Soon Kim

Moments capture a huge part of our lives. Accurate recognition of these moments is challenging due to the diverse and complex interpretation of the moments. Action recognition refers to the act of classifying the desired action/activity…

Computer Vision and Pattern Recognition · Computer Science 2018-09-14 Ankit Shah , Harini Kesavamoorthy , Poorva Rane , Pramati Kalwad , Alexander Hauptmann , Florian Metze

Generalist embodied agents must perform interactive, causally-dependent reasoning, continually interacting with the environment, acquiring information, and updating plans to solve long-horizon tasks before they could be adopted in real-life…

Robotics · Computer Science 2026-04-21 Xianhao Wang , Xiaojian Ma , Haozhe Hu , Rongpeng Su , Yutian Cheng , Zhou Ziheng , Hangxin Liu , Lei Liu , Bin Li , Qing Li

We introduce UCF101 which is currently the largest dataset of human actions. It consists of 101 action classes, over 13k clips and 27 hours of video data. The database consists of realistic user uploaded videos containing camera motion and…

Computer Vision and Pattern Recognition · Computer Science 2012-12-04 Khurram Soomro , Amir Roshan Zamir , Mubarak Shah

We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible research in web agents. It contains 31,725 trajectories and 318k steps, featuring a core…

Artificial Intelligence · Computer Science 2026-04-15 Sicheng Fan , Rui Wan , Yifei Leng , Gaoning Liang , Li Ling , Yanyi Shang , Dehan Kong

While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-rounded common-sense reasoning akin to human visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Kate Sanders , Benjamin Van Durme

Recently, dataset condensation has made significant progress in the image domain. Unlike images, videos possess an additional temporal dimension, which harbors considerable redundant information, making condensation even more crucial.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Yang Chen , Sheng Guo , Bo Zheng , Limin Wang

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA).…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zhou Yu , Dejing Xu , Jun Yu , Ting Yu , Zhou Zhao , Yueting Zhuang , Dacheng Tao

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Baoyu Liang , Qile Su , Shoutai Zhu , Yuchen Liang , Chao Tong

We describe the DeepMind Kinetics human action video dataset. The dataset contains 400 human action classes, with at least 400 video clips for each action. Each clip lasts around 10s and is taken from a different YouTube video. The actions…

We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Yuan Zang , Hao Tan , Seunghyun Yoon , Franck Dernoncourt , Jiuxiang Gu , Kushal Kafle , Chen Sun , Trung Bui

In traffic engineering, vehicle detectors are trained on limited datasets resulting in poor accuracy when deployed in real world applications. Annotating large-scale high quality datasets is challenging. Typically, these datasets have…

Computer Vision and Pattern Recognition · Computer Science 2015-10-08 Justin A. Eichel , Akshaya Mishra , Nicholas Miller , Nicholas Jankovic , Mohan A. Thomas , Tyler Abbott , Douglas Swanson , Joel Keller

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video chain-of-thought (CoT), and instruction tuning on videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Yan Wang , Yawen Zeng , Jingsheng Zheng , Xiaofen Xing , Jin Xu , Xiangmin Xu

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Fadime Sener , Dibyadip Chatterjee , Daniel Shelepov , Kun He , Dipika Singhania , Robert Wang , Angela Yao

Recent multimodal large language models (MLLMs) have advanced video understanding, yet most still "think about videos" ie once a video is encoded, reasoning unfolds entirely in text, treating visual input as a static context. This passive…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Hanoona Rasheed , Mohammed Zumri , Muhammad Maaz , Ming-Hsuan Yang , Fahad Shahbaz Khan , Salman Khan

We present BASKET, a large-scale basketball video dataset for fine-grained skill estimation. BASKET contains 4,477 hours of video capturing 32,232 basketball players from all over the world. Compared to prior skill estimation datasets, our…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yulu Pan , Ce Zhang , Gedas Bertasius