English
Related papers

Related papers: Procedure-Aware Pretraining for Instructional Vide…

200 papers

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Public Code Review (PCR) is developed in the Software Question Answering (SQA) community, assisting developers in exploring high-quality and efficient review services. Current methods on PCR mainly focus on the reviewer's perspective,…

Software Engineering · Computer Science 2025-11-11 Lin Li , Xinchun Yu , Xinyu Chen , Peng Liang

Procedural activities, ranging from routine cooking to complex surgical operations, are highly structured sequences of actions performed in a specific temporal order. Despite the success of current self-supervised learning (SSL) methods on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Chengan Che , Chao Wang , Xinyue Chen , Sophia Tsoka , Luis C. Garcia-Peraza-Herrera

We present a simplified, task-agnostic multi-modal pre-training approach that can accept either video or text input, or both for a variety of end tasks. Existing pre-training are task-specific by adopting either a single cross-modal encoder…

Computer Vision and Pattern Recognition · Computer Science 2021-10-04 Hu Xu , Gargi Ghosh , Po-Yao Huang , Prahal Arora , Masoumeh Aminzadeh , Christoph Feichtenhofer , Florian Metze , Luke Zettlemoyer

Static image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In contrast, abundant…

Computer Vision and Pattern Recognition · Computer Science 2019-12-03 Yiyi Zhang , Li Niu , Ziqi Pan , Meichao Luo , Jianfu Zhang , Dawei Cheng , Liqing Zhang

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Nikita Dvornik , Isma Hadji , Ran Zhang , Konstantinos G. Derpanis , Animesh Garg , Richard P. Wildes , Allan D. Jepson

Understanding web instructional videos is an essential branch of video understanding in two aspects. First, most existing video methods focus on short-term actions for a-few-second-long video clips; these methods are not directly applicable…

Computer Vision and Pattern Recognition · Computer Science 2018-12-07 Shaojie Wang , Wentian Zhao , Ziyi Kou , Chenliang Xu

In this paper, we study the problem of procedure planning in instructional videos, which aims to make a plan (i.e. a sequence of actions) given the current visual observation and the desired goal. Previous works cast this as a sequence…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 Hanlin Wang , Yilu Wu , Sheng Guo , Limin Wang

This paper focuses on task recognition and action segmentation in weakly-labeled instructional videos, where only the ordered sequence of video-level actions is available during training. We propose a two-stream framework, which exploits…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Reza Ghoddoosian , Saif Sayed , Vassilis Athitsos

We are now witnessing significant progress of deep learning methods in a variety of tasks (or datasets) of proteins. However, there is a lack of a standard benchmark to evaluate the performance of different methods, which hinders the…

Machine Learning · Computer Science 2022-09-20 Minghao Xu , Zuobai Zhang , Jiarui Lu , Zhaocheng Zhu , Yangtian Zhang , Chang Ma , Runcheng Liu , Jian Tang

Learning new skills by observing humans' behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans' decision-making processes, focusing on learning a model to plan goal-directed actions in…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Jing Bi , Jiebo Luo , Chenliang Xu

Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion models on ID-matched data achieves state-of-the-art results…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiayi Gao , Changcheng Hua , Qingchao Chen , Yuxin Peng , Yang Liu

Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimodal personalization methods are largely limited to static…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yuanhong Zheng , Ruichuan An , Xiaopeng Lin , Yuxing Liu , Sihan Yang , Huanyu Zhang , Haodong Li , Qintong Zhang , Renrui Zhang , Guopeng Li , Yifan Zhang , Yuheng Li , Wentao Zhang

Procedural text describes dynamic state changes during a step-by-step natural process (e.g., photosynthesis). In this work, we focus on the task of procedural text understanding, which aims to comprehend such documents and track entities'…

Computation and Language · Computer Science 2021-02-16 Zhihan Zhang , Xiubo Geng , Tao Qin , Yunfang Wu , Daxin Jiang

Precisely naming the action depicted in a video can be a challenging and oftentimes ambiguous task. In contrast to object instances represented as nouns (e.g. dog, cat, chair, etc.), in the case of actions, human annotators typically lack a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Kiyoon Kim , Davide Moltisanti , Oisin Mac Aodha , Laura Sevilla-Lara

A knowledge graph (KG) consists of a set of interconnected typed entities and their attributes. Recently, KGs are popularly used as the auxiliary information to enable more accurate, explainable, and diverse user preference recommendations.…

Information Retrieval · Computer Science 2022-04-19 Yuntao Du , Xinjun Zhu , Lu Chen , Ziquan Fang , Yunjun Gao

Video understanding relies on perceiving the global content and modeling its internal connections (e.g., causality, movement, and spatio-temporal correspondence). To learn these interactions, we apply a mask-then-predict pre-training task…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Hao Tan , Jie Lei , Thomas Wolf , Mohit Bansal

Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and cannot evaluate whether models reason about how objects evolve toward task completion. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Wenliang Guo , Yu Kong

The goal of the YouMakeup VQA Challenge 2020 is to provide a common benchmark for fine-grained action understanding in domain-specific videos e.g. makeup instructional videos. We propose two novel question-answering tasks to evaluate…

Computer Vision and Pattern Recognition · Computer Science 2020-04-14 Shizhe Chen , Weiying Wang , Ludan Ruan , Linli Yao , Qin Jin

Accurate video understanding involves reasoning about the relationships between actors, objects and their environment, often over long temporal intervals. In this paper, we propose a message passing graph neural network that explicitly…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Anurag Arnab , Chen Sun , Cordelia Schmid
‹ Prev 1 4 5 6 7 8 10 Next ›