English
Related papers

Related papers: Rekall: Specifying Video Events using Compositions…

200 papers

Multimedia data is highly expressive and has traditionally been very difficult for a machine to interpret. Middleware systems such as complex event processing (CEP) mine patterns from data streams and send notifications to users in a timely…

Artificial Intelligence · Computer Science 2020-10-01 Piyush Yadav , Edward Curry

Rapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks. However, existing video-language models often overlook precise temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Shimin Chen , Xiaohan Lan , Yitian Yuan , Zequn Jie , Lin Ma

Data videos are a powerful medium for visual data based storytelling, combining animated, chart-centric visualizations with synchronized narration. Widely used in journalism, education, and public communication, they help audiences…

Artificial Intelligence · Computer Science 2026-04-29 Ridwan Mahbub , Syem Aziz , Mahir Ahmed , Shadikur Rahman , Mizanur Rahman , Shafiq Joty , Enamul Hoque

Multimodal event argument role labeling (EARL), a task that assigns a role for each event participant (object) in an image is a complex challenge. It requires reasoning over the entire image, the depicted event, and the interactions between…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Hritik Bansal , Po-Nien Kung , P. Jeffrey Brantingham , Kai-Wei Chang , Nanyun Peng

Videos are a commonly-used type of content in learning during Web search. Many e-learning platforms provide quality content, but sometimes educational videos are long and cover many topics. Humans are good in extracting important sections…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Junaid Ahmed Ghauri , Sherzod Hakimov , Ralph Ewerth

Active learning (AL) can reduce annotation costs in surgical video analysis while maintaining model performance. However, traditional AL methods, developed for images or short video clips, are suboptimal for surgical step recognition due to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Nisarg A. Shah , Bardia Safaei , Shameema Sikder , S. Swaroop Vedula , Vishal M. Patel

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Hahyeon Choi , Junhoo Lee , Nojun Kwak

The event-based Vision-Language Model (VLM) recently has made good progress for practical vision tasks. However, most of these works just utilize CLIP for focusing on traditional perception tasks, which obstruct model understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Pengteng Li , Yunfan Lu , Pinghao Song , Wuyang Li , Huizai Yao , Hui Xiong

Event-Level Video Question Answering (EVQA) requires complex reasoning across video events to obtain the visual information needed to provide optimal answers. However, despite significant progress in model performance, few studies have…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Chenyang Lyu , Tianbo Ji , Yvette Graham , Jennifer Foster

Video anomaly detection is of critical practical importance to a variety of real applications because it allows human attention to be focused on events that are likely to be of interest, in spite of an otherwise overwhelming volume of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-17 Guansong Pang , Cheng Yan , Chunhua Shen , Anton van den Hengel , Xiao Bai

Large-scale annotated datasets allow AI systems to learn from and build upon the knowledge of the crowd. Many crowdsourcing techniques have been developed for collecting image annotations. These techniques often implicitly rely on the fact…

Human-Computer Interaction · Computer Science 2016-10-07 Gunnar A. Sigurdsson , Olga Russakovsky , Ali Farhadi , Ivan Laptev , Abhinav Gupta

Multimodal ML models can process data in multiple modalities (e.g., video, images, audio, text) and are useful for video content analysis in a variety of problems (e.g., object detection, scene understanding). In this paper, we focus on the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Palash Goyal , Saurabh Sahu , Shalini Ghosh , Chul Lee

Disciplines such as business process management and process mining aid organizations by discovering insights about processes on the basis of recorded event data. However, an obstacle to process analysis is data multi-modality: for instance,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Marco Pegoraro , Jonas Seng , Dustin Heller , Wil M. P. van der Aalst , Kristian Kersting

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are…

Computer Vision and Pattern Recognition · Computer Science 2015-05-05 Vignesh Ramanathan , Kevin Tang , Greg Mori , Li Fei-Fei

Video Large Multimodal Models (VLMMs) have shown impressive performance in video understanding, yet their ability to accurately capture the temporal order of multiple events remains underexplored. We interestingly observe that, even when…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Daechul Ahn , Yura Choi , Hyeonbeom Choi , Seongwon Cho , San Kim , Jonghyun Choi

This paper investigates the challenge of extracting highlight moments from videos. To perform this task, we need to understand what constitutes a highlight for arbitrary video domains while at the same time being able to scale across…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 David Chuan-En Lin , Fabian Caba Heilbron , Joon-Young Lee , Oliver Wang , Nikolas Martelaro

Continual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced sequentially. For video data, the task becomes more complex…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Tieyuan Chen , Huabin Liu , Chern Hong Lim , John See , Xing Gao , Junhui Hou , Weiyao Lin

Content metadata plays a very important role in movie recommender systems as it provides valuable information about various aspects of a movie such as genre, cast, plot synopsis, box office summary, etc. Analyzing the metadata can help…

Information Retrieval · Computer Science 2023-09-19 Saurabh Agrawal , John Trenkle , Jaya Kawale

We address the problem of video representation learning without human-annotated labels. While previous efforts address the problem by designing novel self-supervised tasks using video data, the learned features are merely on a…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Jiangliu Wang , Jianbo Jiao , Linchao Bao , Shengfeng He , Yunhui Liu , Wei Liu

Understanding temporal information and how the visual world changes over time is a fundamental ability of intelligent systems. In video understanding, temporal information is at the core of many current challenges, including compression,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-31 Laura Sevilla-Lara , Shengxin Zha , Zhicheng Yan , Vedanuj Goswami , Matt Feiszli , Lorenzo Torresani
‹ Prev 1 8 9 10 Next ›