English
Related papers

Related papers: STREAM: Spatio-TempoRal Evaluation and Analysis Me…

200 papers

Event-based vision, characterized by low redundancy, focus on dynamic motion, and inherent privacy-preserving properties, naturally fits the demands of video anomaly detection (VAD). However, the absence of dedicated event-stream anomaly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Peng Wu , Yuting Yan , Guansong Pang , Yujia Sun , Qingsen Yan , Peng Wang , Yanning Zhang

Current text-to-video models (T2V) can generate high-quality, temporally coherent, and visually realistic videos. Nonetheless, errors still often occur, and are more nuanced and local compared to the previous generation of T2V models. While…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Aditya Chinchure , Sahithya Ravi , Pushkar Shukla , Vered Shwartz , Leonid Sigal

Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require streaming perception and proactive interaction, where video…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Junho Kim , Hosu Lee , James M. Rehg , Minsu Kim , Yong Man Ro

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junyoung Seo , Jisang Han , Jaewoo Jung , Siyoon Jin , Joungbin Lee , Takuya Narihira , Kazumi Fukuda , Takashi Shibuya , Donghoon Ahn , Shoukang Hu , Seungryong Kim , Yuki Mitsufuji

Diffusion generative models have recently become a powerful technique for creating and modifying high-quality, coherent video content. This survey provides a comprehensive overview of the critical components of diffusion models for video…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Andrew Melnik , Michal Ljubljanac , Cong Lu , Qi Yan , Weiming Ren , Helge Ritter

Understanding long videos with multimodal large language models (MLLMs) remains challenging due to the heavy redundancy across frames and the need for temporally coherent representations. Existing static strategies, such as sparse sampling,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Naishan Zheng , Jie Huang , Qingpei Guo , Feng Zhao

Volumetric video streaming offers immersive 3D experiences but faces significant challenges due to high bandwidth requirements and latency issues in transmitting detailed content in real time. Traditional methods like point cloud streaming…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Boyan Li , Yongting Chen , Dayou Zhang , Fangxin Wang

As with many machine learning problems, the progress of image generation methods hinges on good evaluation metrics. One of the most popular is the Frechet Inception Distance (FID). FID estimates the distance between a distribution of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-29 Sadeep Jayasumana , Srikumar Ramalingam , Andreas Veit , Daniel Glasner , Ayan Chakrabarti , Sanjiv Kumar

Video generation is an inherently challenging task, as it requires modeling realistic temporal dynamics as well as spatial content. Existing methods entangle the two intrinsically different tasks of motion and content creation in a single…

Computer Vision and Pattern Recognition · Computer Science 2020-01-13 Ximeng Sun , Huijuan Xu , Kate Saenko

Large Language Models have shown remarkable efficacy in generating streaming data such as text and audio, thanks to their temporally uni-directional attention mechanism, which models correlations between the current token and previous…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Zhening Xing , Gereon Fox , Yanhong Zeng , Xingang Pan , Mohamed Elgharib , Christian Theobalt , Kai Chen

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Shaobin Zhuang , Zhipeng Huang , Binxin Yang , Ying Zhang , Fangyikang Wang , Canmiao Fu , Chong Sun , Zheng-Jun Zha , Chen Li , Yali Wang

Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensitivity of conventional cameras. In this work, we explore gait recognition using event…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Senyan Xu , Shuai Chen , Chuanfu Shen , Kean Liu , Zhijing Sun , Chengzhi Cao , Xueyang Fu

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yanxiao Sun , Jiafu Wu , Yun Cao , Chengming Xu , Yabiao Wang , Weijian Cao , Donghao Luo , Chengjie Wang , Yanwei Fu

Video streaming is a fundamental Internet service, while the quality still cannot be guaranteed especially in poor network conditions such as bandwidth-constrained and remote areas. Existing works mainly work towards two directions:…

Networking and Internet Architecture · Computer Science 2026-02-04 Tianyi Gong , Zijian Cao , Zixing Zhang , Jiangkai Wu , Xinggong Zhang , Shuguang Cui , Fangxin Wang

This thesis is part of a CIFRE agreement between the company Othello and the LIASD laboratory. The objective is to develop an artificial intelligence system that can detect real-time dangers in a video stream. To achieve this, a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Fabien Poirier

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Xavier Thomas , Youngsun Lim , Ananya Srinivasan , Audrey Zheng , Deepti Ghadiyaram

Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet remains difficult due to the lack of suitable data and deployable editing models.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yiren Song , Cheng Liu , Yuxin Jiang , Mike Zheng Shou

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains limited due to storage constraints of historical visual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xiangyu Zeng , Kefan Qiu , Qingyu Zhang , Xinhao Li , Jing Wang , Jiaxin Li , Ziang Yan , Kun Tian , Meng Tian , Xinhai Zhao , Yi Wang , Limin Wang

Online Video Large Language Models (VideoLLMs) play a critical role in supporting responsive, real-time interaction. Existing methods focus on streaming perception, lacking a synchronized logical reasoning stream. However, directly applying…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yiran Guan , Liang Yin , Dingkang Liang , Jianzhong Ju , Zhenbo Luo , Jian Luan , Yuliang Liu , Xiang Bai

Immersive video offers the freedom to navigate inside virtualized environment. Instead of streaming the bulky immersive videos entirely, a viewport (also referred to as field of view, FoV) adaptive streaming is preferred. We often stream…

Multimedia · Computer Science 2018-02-19 Shaowei Xie , Qiu Shen , Yiling Xu , Qiaojian Qian , Shaowei Wang , Zhan Ma , Wenjun Zhang