English
Related papers

Related papers: StreamPro: From Reactive Perception to Proactive D…

200 papers

In gradient-based learning, a step size chosen in parameter units does not produce a predictable per-step change in function output. This often leads to instability in the streaming setting (i.e., batch size=1), where stochasticity is not…

Machine Learning · Computer Science 2026-04-22 Arsalan Sharifnassab , Mohamed Elsayed , Kris De Asis , A. Rupam Mahmood , Richard S. Sutton

Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-free framework that enables a frozen video generator to assess and improve the physical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Mariam Hassan , Kaouther Messaoud , Wuyang Li , Alexandre Alahi

Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all change over time. Static bias benchmarks remain useful, but they do not show how models…

Computation and Language · Computer Science 2026-05-29 Mohd Ariful Haque , Fahad Rahman , Kishor Datta Gupta , Roy George

Multimodal Large Language Models excel at offline audio-visual understanding, but their ability to serve as mobile assistants in continuous real-world streams remains underexplored. In daily phone use, mobile assistants must track streaming…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Xudong Lu , Huankang Guan , Yang Bo , Jinpeng Chen , Xintong Guo , Shuhan Li , Fang Liu , Peiwen Sun , Xueying Li , Wei Zhang , Xue Yang , Rui Liu , Hongsheng Li

The Quality of Experience (QoE) of streaming service is often degraded by frequent playback interruptions. To mitigate the interruptions, the media player prefetches streaming contents before starting playback, at a cost of delay. We study…

Networking and Internet Architecture · Computer Science 2014-06-06 Yuedong Xu , Salaheddine Elayoubi , Eitan Altman , Rachid El-Azouzi , Yinghao Yu

Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge model sizes, limiting their real-time applicability. We…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Xiao Yu , Yan Fang , Xiaojie Jin , Yao Zhao , Yunchao Wei

The emerging field of action prediction plays a vital role in various computer vision applications such as autonomous driving, activity analysis and human-computer interaction. Despite significant advancements, accurately predicting future…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Izzeddin Teeti , Rongali Sai Bhargav , Vivek Singh , Andrew Bradley , Biplab Banerjee , Fabio Cuzzolin

The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity. Addressing this gap, we introduce the Real-tIme Video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Yansong Shi , Qingsong Zhao , Tianxiang Jiang , Xiangyu Zeng , Yi Wang , Limin Wang

Evaluation of recommender systems is typically done with finite datasets. This means that conventional evaluation methodologies are only applicable in offline experiments, where data and models are stationary. However, in real world…

Information Retrieval · Computer Science 2015-05-04 João Vinagre , Alípio Mário Jorge , João Gama

With the rapid growth of video centered social media, the ability to anticipate risky events from visual data is a promising direction for ensuring public safety and preventing real world accidents. Prior work has extensively studied…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Sha Luo , Yogesh Prabhu , Timothy Ossowski , Kaiping Chen , Junjie Hu

Human action recognition in video is an active yet challenging research topic due to high variation and complexity of data. In this paper, a novel video based action recognition framework utilizing complementary cues is proposed to handle…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Muhammad Usman Khalid , Jie Yu

In this work, we investigate diffusion-based video prediction models, which forecast future video frames, for continuous video streams. In this context, the models observe continuously new training samples, and we aim to leverage this to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Sina Mokhtarzadeh Azar , Emad Bahrami , Enrico Pallotta , Gianpiero Francesca , Radu Timofte , Juergen Gall

With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce StreamMind, a video LLM framework that achieves ultra-FPS streaming…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Xin Ding , Hao Wu , Yifan Yang , Shiqi Jiang , Donglin Bai , Zhibo Chen , Ting Cao

Action recognition is a key problem in computer vision that labels videos with a set of predefined actions. Capturing both, semantic content and motion, along the video frames is key to achieve high accuracy performance on this task. Most…

Computer Vision and Pattern Recognition · Computer Science 2019-10-23 Xia Huang , Hossein Mousavi , Gemma Roig

Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Tinghui Zhu , Sheng Zhang , James Y. Huang , Selena Song , Xiaofei Wen , Yuankai Li , Hoifung Poon , Muhao Chen

We present the Streaming Reservoir Convergence Theorem (SRCT), a novel mathematical framework for multi-provider adaptive bitrate streaming that addresses three fundamental structural weaknesses in current systems: linear provider probing,…

Gesture recognition in resource-constrained scenarios faces significant challenges in achieving high accuracy and low latency. The streaming gesture recognition framework, Duo Streamers, proposed in this paper, addresses these challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Boxuan Zhu , Sicheng Yang , Zhuo Wang , Haining Liang , Junxiao Shen

Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recognition, use a video backbone to directly extract…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Yucheng Zhao , Chong Luo , Chuanxin Tang , Dongdong Chen , Noel Codella , Zheng-Jun Zha

Recently, HTTP-Based Adaptive Streaming has become the de facto standard for video streaming over the Internet. It allows the client to adapt media characteristics to varying network conditions in order to maximize Quality of Experience…

Networking and Internet Architecture · Computer Science 2016-03-04 Konstantin Miller , Abdel-Karim Al-Tamimi , Adam Wolisz

In modern human-robot collaboration (HRC) applications, multiple perception modules jointly extract visual, auditory, and contextual cues to achieve comprehensive scene understanding, enabling the robot to provide appropriate assistance to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Dingcheng Huang , Xiaotong Zhang , Kamal Youcef-Toumi
‹ Prev 1 8 9 10 Next ›