English
Related papers

Related papers: Veritas: Answering Causal Queries from Video Strea…

200 papers

The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce VideoVeritas, a framework that integrates fine-grained perception and fact-based…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Hao Tan , Jun Lan , Senyuan Shi , Zichang Tan , Zijian Yu , Huijia Zhu , Weiqiang Wang , Jun Wan , Zhen Lei

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs before processing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Xingyi Zhou , Anurag Arnab , Shyamal Buch , Shen Yan , Austin Myers , Xuehan Xiong , Arsha Nagrani , Cordelia Schmid

Streaming data routinely generated by mobile phones, social networks, e-commerce, and electronic health records present new opportunities for near real-time surveillance of the impact of an intervention on an outcome of interest via causal…

Methodology · Statistics 2021-11-30 Xu Shi , Lan Luo

Efficient vision works maximize accuracy under a latency budget. These works evaluate accuracy offline, one image at a time. However, real-time vision applications like autonomous driving operate in streaming settings, where ground truth…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Gur-Eyal Sela , Ionel Gog , Justin Wong , Kumar Krishna Agrawal , Xiangxi Mo , Sukrit Kalra , Peter Schafhalter , Eric Leong , Xin Wang , Bharathan Balaji , Joseph Gonzalez , Ion Stoica

Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline scenario. However, online video streams, as one of the most…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Haoji Zhang , Yiqin Wang , Yansong Tang , Yong Liu , Jiashi Feng , Jifeng Dai , Xiaojie Jin

Deep neural networks facilitate video question answering (VideoQA), but the real-world applications on video streams such as CCTV and live cast place higher demands on the solver. To address the challenges of VideoQA on long videos of…

Multimedia · Computer Science 2023-03-08 Weikai Kong , Shuhong Ye , Chenglin Yao , Jianfeng Ren

The increase in video streaming has presented a challenge of handling stream request effectively, especially over networks that are variable. This paper describes a new adaptive video streaming architecture capable of changing the video…

Networking and Internet Architecture · Computer Science 2025-02-05 Mohammad Tarik , Qutaiba Ibrahim

Being able to automatically and quickly understand the user context during a session is a main issue for recommender systems. As a first step toward achieving that goal, we propose a model that observes in real time the diversity brought by…

Information Retrieval · Computer Science 2016-01-11 Sylvain Castagnos , Amaury L 'Huillier , Anne Boyer

Although the problem of automatic video summarization has recently received a lot of attention, the problem of creating a video summary that also highlights elements relevant to a search query has been less studied. We address this problem…

Computer Vision and Pattern Recognition · Computer Science 2017-09-29 Arun Balajee Vasudevan , Michael Gygli , Anna Volokitin , Luc Van Gool

Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Dong Zhuo , Wenzhao Zheng , Jiahe Guo , Yuqi Wu , Jie Zhou , Jiwen Lu

Motivated by emerging vision-based intelligent services, we consider the problem of rate adaptation for high quality and low delay visual information delivery over wireless networks using scalable video coding. Rate adaptation in this…

Multimedia · Computer Science 2017-04-11 Hussein Al-Zubaidy , Viktoria Fodor , György Dán , Markus Flierl

Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall short due to their implicit, query-agnostic modeling of visual evidence. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Ke Ma , Jiaqi Tang , Bin Guo , Xueting Han , Ruonan Xu , Qingfeng He , Ziheng Wang , Xu Wang , Qifeng Chen , Zhiwen Yu , Yunhao Liu

HTTP-based video streaming technologies allow for flexible rate selection strategies that account for time-varying network conditions. Such rate changes may adversely affect the user's Quality of Experience; hence online prediction of the…

Multimedia · Computer Science 2017-07-11 Christos G. Bampis , Alan C. Bovik

Experience of live video streaming can be improved if the video uploader has more accurate knowledge about the future available bandwidth. Because with such knowledge, one is able to know what sizes should he encode the frames to be in an…

Multimedia · Computer Science 2022-10-05 Weijia Zheng

Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Dibyadip Chatterjee , Zhanzhong Pang , Fadime Sener , Yale Song , Angela Yao

This paper introduces a scalable causal inference framework for estimating the immediate, session-level effects of on-demand human tutoring embedded within adaptive learning systems. Because students seek assistance at moments of…

Human-Computer Interaction · Computer Science 2026-02-24 Kirk Vanacore , Danielle R Thomas , Digory Smith , Bibi Groot , Justin Reich , Rene Kizilcec

Given an input video of a person and a new garment, the objective of this paper is to synthesize a new video where the person is wearing the specified garment while maintaining spatiotemporal consistency. Although significant advances have…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Hung Nguyen , Quang Qui-Vinh Nguyen , Khoi Nguyen , Rang Nguyen

Video streaming analytics is a crucial workload for vision-language model serving, but the high cost of multimodal inference limits scalability. Prior systems reduce inference cost by exploiting temporal and spatial redundancy in video…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-10 Yulin Zou , Yan Chen , Wenyan Chen , JooYoung Park , Shivaraman Nitin , Luo Tao , Francisco Romero , Dmitrii Ustiugov

VERSA provides a general-purpose framework for defining and recognizing events in live or recorded surveillance video streams. The approach for event recognition in VERSA is using a declarative logic language to define the spatial and…

Computer Vision and Pattern Recognition · Computer Science 2010-07-23 Stephen O'Hara

Understanding and tuning the performance of extreme-scale parallel computing systems demands a streaming approach due to the computational cost of applying offline algorithms to vast amounts of performance log data. Analyzing large…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-28 Suraj P. Kesavan , Takanori Fujiwara , Jianping Kelvin Li , Caitlin Ross , Misbah Mubarak , Christopher D. Carothers , Robert B. Ross , Kwan-Liu Ma