English
Related papers

Related papers: SPIKE-RL: Video-LLMs meet Bayesian Surprise

200 papers

Video summarization helps turn long videos into clear, concise representations that are easier to review, document, and analyze, especially in high-stakes domains like surgical training. Prior work has progressed from using basic visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Shreya Rajpal , Michal Golovanevsky , Carsten Eickhoff

Tracking the dynamics of non-canonical biological systems in microscopy videos remains a persistent challenge. Both classical and learning-based trackers depend on expert-reviewed data to be evaluated and adapted, yet exhaustive manual…

Procedural activities, ranging from routine cooking to complex surgical operations, are highly structured sequences of actions performed in a specific temporal order. Despite the success of current self-supervised learning (SSL) methods on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Chengan Che , Chao Wang , Xinyue Chen , Sophia Tsoka , Luis C. Garcia-Peraza-Herrera

The emergence of large pre-trained vision-language models (VLMs) represents a paradigm shift in machine learning, with unprecedented results in a broad span of visual recognition tasks. CLIP, one of the most popular VLMs, has exhibited…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Pablo Morales-Álvarez , Stergios Christodoulidis , Maria Vakalopoulou , Pablo Piantanida , Jose Dolz

Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Dibyadip Chatterjee , Zhanzhong Pang , Fadime Sener , Yale Song , Angela Yao

This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, surprisingly, they often do not. We introduce REVEAL{}, a…

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Zichen Liu , Kunlun Xu , Bing Su , Xu Zou , Yuxin Peng , Jiahuan Zhou

Large vision-language contrastive models (VLCMs), such as CLIP, have become foundational, demonstrating remarkable success across a variety of downstream tasks. Despite their advantages, these models, akin to other foundational systems,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Haocheng Dai , Sarang Joshi

Video summarization methods are usually classified into shot-level or frame-level methods, which are individually used in a general way. This paper investigates the underlying complementarity between the frame-level and shot-level methods,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Yubo An , Shenghui Zhao , Guoqiang Zhang

Video processing solutions for motion analysis are key tasks in many computer vision applications, ranging from human activity recognition to object detection. In particular, speed estimation algorithms may be relevant in contexts such as…

Image and Video Processing · Electrical Eng. & Systems 2022-11-29 Veronica Mattioli , Davide Alinovi , Riccardo Raheli

As a bio-inspired vision sensor, the spike camera emulates the operational principles of the fovea, a compact retinal region, by employing spike discharges to encode the accumulation of per-pixel luminance intensity. Leveraging its high…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Lin Zhu , Xianzhang Chen , Xiao Wang , Hua Huang

We evaluate two popular local explainability techniques, LIME and SHAP, on a movie recommendation task. We discover that the two methods behave very differently depending on the sparsity of the data set. LIME does better than SHAP in dense…

Machine Learning · Computer Science 2022-06-13 Claudia V. Roberts , Ehtsham Elahi , Ashok Chandrashekar

Long-form video reasoning remains a major challenge for Video Large Language Models (Video LLMs), as static uniform frame sampling leads to information dilution and obscures critical evidence. Furthermore, existing pixel-space video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Xuchen Li , Xuzhao Li , Shiyu Hu , Kaiqi Huang

While Video Large Language Models (Video-LLMs) have demonstrated remarkable performance across general video understanding benchmarks-particularly in video captioning and descriptive tasks-they consistently underperform on tasks that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Sameep Vani , Shreyas Jena , Maitreya Patel , Chitta Baral , Somak Aditya , Yezhou Yang

Deep learning has led to remarkable advances in computer vision. Even so, today's best models are brittle when presented with variations that differ even slightly from those seen during training. Minor shifts in the pose, color, or…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Mark Ibrahim , Diane Bouchacourt , Ari Morcos

Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with varying camera angles or scene changes. This challenge can…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yujia Liang , Jile Jiao , Xuetao Feng , Zixuan Ye , Yuan Wang , Zhicheng Wang

Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI tutors and human educators. Yet such simulators are typically evaluated by output…

Computation and Language · Computer Science 2026-05-14 Heejin Do , Shashank Sonkar , Mrinmaya Sachan

Despite the recent advances of the artificial intelligence, building social intelligence remains a challenge. Among social signals, laughter is one of the distinctive expressions that occurs during social interactions between humans. In…

Computation and Language · Computer Science 2024-05-27 Lee Hyun , Kim Sung-Bin , Seungju Han , Youngjae Yu , Tae-Hyun Oh

Videos are a popular media form, where online video streaming has recently gathered much popularity. In this work, we propose a novel method of real-time video stabilization - transforming a shaky video to a stabilized video as if it were…

Computer Vision and Pattern Recognition · Computer Science 2021-11-12 Jinsoo Choi , Jaesik Park , In So Kweon

Non-Intrusive Load Monitoring (NILM) is a field of research focused on segregating constituent electrical loads in a system based only on their aggregated signal. Significant computational resources and research time are spent training…

Artificial Intelligence · Computer Science 2020-09-17 Richard Jones , Christoph Klemenjak , Stephen Makonin , Ivan V. Bajic
‹ Prev 1 4 5 6 7 8 10 Next ›