English
Related papers

Related papers: Induce, Edit, Retrieve: Language Grounded Multimod…

200 papers

We address the problem of text-based activity retrieval in video. Given a sentence describing an activity, our task is to retrieve matching clips from an untrimmed video. To capture the inherent structures present in both text and video, we…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Huijuan Xu , Kun He , Bryan A. Plummer , Leonid Sigal , Stan Sclaroff , Kate Saenko

Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encoding an entire clip or video segment. However, queries often describe an action or event over…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Quoc-Bao Nguyen-Le , Thanh-Huy Le-Nguyen

This paper presents a novel retrieval pipeline for video collections, which aims to retrieve the most significant parts of an edited video for a given query, and represent them with thumbnails which are at the same time semantically…

Computer Vision and Pattern Recognition · Computer Science 2016-04-12 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

Human communication takes many forms, including speech, text and instructional videos. It typically has an underlying structure, with a starting point, ending, and certain objective steps between them. In this paper, we consider…

Computer Vision and Pattern Recognition · Computer Science 2016-05-12 Ozan Sener , Amir Roshan Zamir , Chenxia Wu , Silvio Savarese , Ashutosh Saxena

Expertise is often built by learning from examples. This process, known as schema induction, helps us identify patterns from examples. Despite its importance, schema induction remains a challenging cognitive task. Recent advances in…

Human-Computer Interaction · Computer Science 2025-02-24 Sitong Wang , Lydia B. Chilton

Online resources such as WikiHow compile a wide range of scripts for performing everyday tasks, which can assist models in learning to reason about procedures. However, the scripts are always presented in a linear manner, which does not…

Computation and Language · Computer Science 2023-05-30 Yu Zhou , Sha Li , Manling Li , Xudong Lin , Shih-Fu Chang , Mohit Bansal , Heng Ji

Procedural activity understanding requires perceiving human actions in terms of a broader task, where multiple keysteps are performed in sequence across a long video to reach a final goal state -- such as the steps of a recipe or a DIY…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Kumar Ashutosh , Santhosh Kumar Ramakrishnan , Triantafyllos Afouras , Kristen Grauman

Human communication typically has an underlying structure. This is reflected in the fact that in many user generated videos, a starting point, ending, and certain objective steps between these two can be identified. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2016-01-28 Ozan Sener , Amir Zamir , Silvio Savarese , Ashutosh Saxena

Computational Video Editing Systems output video generally follows a particular form, e.g. conversation or music videos, in this way they are domain specific. We describe a recent development in our video annotation and segmentation system…

Multimedia · Computer Science 2021-02-23 Sean Butler

Creative and communicative work is often underpinned by implicit structures, such as the Hero's Journey in storytelling, design patterns in software, or chord progressions in music. People often learn these structures from examples - a…

Human-Computer Interaction · Computer Science 2026-04-10 Sitong Wang , Samia Menon , Dingzeyu Li , Xiaojuan Ma , Richard Zemel , Lydia B. Chilton

This paper presents an integrated multi-agents architecture for indexing and retrieving video information.The focus of our work is to elaborate an extensible approach that gathers a priori almost of the mandatory tools which palliate to the…

Information Retrieval · Computer Science 2014-08-01 Yasser El Madani El Alami , El Habib Nfaoui , Omar El Beqqali

Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text. While previous work focuses on building systems for inducing grammars on text that are well-aligned with…

Computation and Language · Computer Science 2022-10-25 Songyang Zhang , Linfeng Song , Lifeng Jin , Haitao Mi , Kun Xu , Dong Yu , Jiebo Luo

Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 G. Thomas Hudson , Dean Slack , Thomas Winterbottom , Jamie Sterling , Chenghao Xiao , Junjie Shentu , Noura Al Moubayed

When obtaining visual illustrations from text descriptions, today's methods take a description with a single text context - a caption, or an action description - and retrieve or generate the matching visual context. However, prior work does…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Chi Hsuan Wu , Kumar Ashutosh , Kristen Grauman

We address the problem of automatically learning the main steps to complete a certain task, such as changing a car tire, from a set of narrated instruction videos. The contributions of this paper are three-fold. First, we develop a new…

Computer Vision and Pattern Recognition · Computer Science 2016-06-29 Jean-Baptiste Alayrac , Piotr Bojanowski , Nishant Agrawal , Josef Sivic , Ivan Laptev , Simon Lacoste-Julien

In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Panwen Hu , Nan Xiao , Feifei Li , Yongquan Chen , Rui Huang

With the rapid advancement of multimodal information retrieval, increasingly complex retrieval tasks have emerged. Existing methods predominately rely on task-specific fine-tuning of vision-language models, often those trained with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Yikun Liu , Pingan Chen , Jiayin Cai , Xiaolong Jiang , Yao Hu , Jiangchao Yao , Yanfeng Wang , Weidi Xie

Though pre-training vision-language models have demonstrated significant benefits in boosting video-text retrieval performance from large-scale web videos, fine-tuning still plays a critical role with manually annotated clips with start and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Bin Zhu , Kevin Flanagan , Adriano Fragomeni , Michael Wray , Dima Damen

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable…

Computer Vision and Pattern Recognition · Computer Science 2016-11-11 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

Schema induction builds a graph representation explaining how events unfold in a scenario. Existing approaches have been based on information retrieval (IR) and information extraction(IE), often with limited human curation. We demonstrate a…

‹ Prev 1 2 3 10 Next ›