English
Related papers

Related papers: GUIDE: A Guideline-Guided Dataset for Instructiona…

200 papers

In this paper we present an approach for localizing steps of procedural activities in narrated how-to videos. To deal with the scarcity of labeled data at scale, we source the step descriptions from a language knowledge base (wikiHow)…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Effrosyni Mavroudi , Triantafyllos Afouras , Lorenzo Torresani

Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives. These procedures serve as a guiding framework that helps to achieve goals efficiently, whether it is…

We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose InstructDoc, the…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Ryota Tanaka , Taichi Iki , Kyosuke Nishida , Kuniko Saito , Jun Suzuki

While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-rounded common-sense reasoning akin to human visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Kate Sanders , Benjamin Van Durme

Human communication takes many forms, including speech, text and instructional videos. It typically has an underlying structure, with a starting point, ending, and certain objective steps between them. In this paper, we consider…

Computer Vision and Pattern Recognition · Computer Science 2016-05-12 Ozan Sener , Amir Roshan Zamir , Chenxia Wu , Silvio Savarese , Ashutosh Saxena

Despite the number of currently available datasets on video question answering, there still remains a need for a dataset involving multi-step and non-factoid answers. Moreover, relying on video transcripts remains an under-explored topic.…

Computation and Language · Computer Science 2020-06-02 Anthony Colas , Seokhwan Kim , Franck Dernoncourt , Siddhesh Gupte , Daisy Zhe Wang , Doo Soon Kim

Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Chi Zhang , Chengjian Feng , Feng Yan , Qiming Zhang , Mingjin Zhang , Yujie Zhong , Jing Zhang , Lin Ma

Recent research in behaviour understanding through language grounding has shown it is possible to automatically generate behaviour models from textual instructions. These models usually have goal-oriented structure and are modelled with…

Artificial Intelligence · Computer Science 2020-01-14 Debajyoti Paul Chowdhury , Arghya Biswas , Tomasz Sosnowski , Kristina Yordanova

People increasingly use videos on the Web as a source for learning. To support this way of learning, researchers and developers are continuously developing tools, proposing guidelines, analyzing data, and conducting experiments. However, it…

Multimedia · Computer Science 2023-08-15 Evelyn Navarrete , Andreas Nehring , Sascha Schanze , Ralph Ewerth , Anett Hoppe

Generating video descriptions in natural language (a.k.a. video captioning) is a more challenging task than image captioning as the videos are intrinsically more complicated than images in two aspects. First, videos cover a broader range of…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Shizhe Chen , Jia Chen , Qin Jin

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Lavisha Aggarwal , Vikas Bahirwani , Andrea Colaco

Deep learning has shown remarkable progress in a wide range of problems. However, efficient training of such models requires large-scale datasets, and getting annotations for such datasets can be challenging and costly. In this work, we…

Multimedia · Computer Science 2021-10-14 Mohit Sharma , Raj Patra , Harshal Desai , Shruti Vyas , Yogesh Rawat , Rajiv Ratn Shah

Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction of high-quality instruction datasets is crucial for…

Artificial Intelligence · Computer Science 2026-02-12 Li Du , Hanyu Zhao , Yiming Ju , Tengfei Pan

Learning from (procedural) videos has increasingly served as a pathway for embodied agents to acquire skills from human demonstrations. To do this, video understanding models must be able to obtain structured understandings, such as the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zitian Tang , Rohan Myer Krishnan , Zhiqiu Yu , Chen Sun

Video description is one of the most challenging problems in vision and language understanding due to the large variability both on the video and language side. Models, hence, typically shortcut the difficulty in recognition and generate…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Luowei Zhou , Yannis Kalantidis , Xinlei Chen , Jason J. Corso , Marcus Rohrbach

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Kumar Ashutosh , Chi Hsuan Wu , Kristen Grauman

Video description involves the generation of the natural language description of actions, events, and objects in the video. There are various applications of video description by filling the gap between languages and vision for visually…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Alok Singh , Thoudam Doren Singh , Sivaji Bandyopadhyay

Many high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e. procedures) that…

Computation and Language · Computer Science 2020-05-20 Angela S. Lin , Sudha Rao , Asli Celikyilmaz , Elnaz Nouri , Chris Brockett , Debadeepta Dey , Bill Dolan

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Ming Hu , Lin Wang , Siyuan Yan , Don Ma , Qingli Ren , Peng Xia , Wei Feng , Peibo Duan , Lie Ju , Zongyuan Ge

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily generalizable to diverse tasks, domains, and datasets. To…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Linjie Li , Jie Lei , Zhe Gan , Licheng Yu , Yen-Chun Chen , Rohit Pillai , Yu Cheng , Luowei Zhou , Xin Eric Wang , William Yang Wang , Tamara Lee Berg , Mohit Bansal , Jingjing Liu , Lijuan Wang , Zicheng Liu