中文
相关论文

相关论文: SPIKE-RL: Video-LLMs meet Bayesian Surprise

200 篇论文

Ensuring transparency in machine learning decisions is critically important, especially in sensitive sectors such as healthcare, finance, and justice. Despite this, some popular explainable algorithms, such as Local Interpretable…

机器学习 · 计算机科学 2025-03-27 Shakiba Rahimiaghdam , Hande Alemdar

The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Jingjun Yang , Liangwei Fan , Jinpu Zhang , Xiangkai Lian , Hui Shen , Dewen Hu

Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards and multi-objective alignment, RMs are expected to encode…

机器学习 · 计算机科学 2026-05-21 Jiwoo Hong , Shao Tang , Zhipeng Wang

Real-time Video Frame Interpolation (VFI) has long been dominated by flow-based methods like RIFE, which offer high throughput but often fail in complicated scenarios involving large motion and occlusion. Conversely, recent diffusion-based…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Pan Ben Wong , Chengli Wu , Hanyue Lu

Dynamic scenes contain intricate spatio-temporal information, crucial for mobile robots, UAVs, and autonomous driving systems to make informed decisions. Parsing these scenes into semantic triplets <Subject-Predicate-Object> for accurate…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Hang Zhang , Zhuoling Li , Jun Liu

Anomaly detection in videos aims at reporting anything that does not conform the normal behaviour or distribution. However, due to the sparsity of abnormal video clips in real life, collecting annotated data for supervised learning is…

计算机视觉与模式识别 · 计算机科学 2019-10-22 Yiwei Lu , Mahesh Kumar Krishna Reddy , Seyed shahabeddin Nabavi , Yang Wang

Recently, memory-based approaches show promising results on semi-supervised video object segmentation. These methods predict object masks frame-by-frame with the help of frequently updated memory of the previous mask. Different from this…

计算机视觉与模式识别 · 计算机科学 2022-08-04 Kwanyong Park , Sanghyun Woo , Seoung Wug Oh , In So Kweon , Joon-Young Lee

We introduce an audiovisual method for long-range text-to-video retrieval. Unlike previous approaches designed for short video retrieval (e.g., 5-15 seconds in duration), our approach aims to retrieve minute-long videos that capture complex…

计算机视觉与模式识别 · 计算机科学 2022-08-03 Yan-Bo Lin , Jie Lei , Mohit Bansal , Gedas Bertasius

Although speculative decoding is widely used to accelerate Vision-Language Models (VLMs) inference, it faces severe performance collapse when applied to Video Large Language Models (Vid-LLMs). The draft model typically falls into the trap…

计算机视觉与模式识别 · 计算机科学 2026-02-18 Libo Zhang , Zhaoning Zhang , Wangyang Hong , Peng Qiao , Dongsheng Li

Prior studies on Video Anomaly Detection (VAD) mainly focus on detecting whether each video frame is abnormal or not in the video, which largely ignore the structured video semantic information (i.e., what, when, and where does the abnormal…

计算机视觉与模式识别 · 计算机科学 2025-02-27 Junxiao Ma , Jingjing Wang , Jiamin Luo , Peiying Yu , Guodong Zhou

We propose a novel framework for abnormal event detection in video that requires no training sequences. Our framework is based on unmasking, a technique previously used for authorship verification in text documents, which we adapt to our…

计算机视觉与模式识别 · 计算机科学 2017-07-26 Radu Tudor Ionescu , Sorina Smeureanu , Bogdan Alexe , Marius Popescu

Large multimodal models (LMMs) have recently demonstrated remarkable performance in video question answering (VideoQA), yet reasoning over video remains challenging due to high inference cost and diluted information. Keyframe selection…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Minchan Kwon , Hyounguk Shon , Junmo Kim

Local explanation methods such as LIME (Ribeiro et al., 2016) remain fundamental to trustworthy AI, yet their application to NLP is limited by a reliance on random token masking. These heuristic perturbations frequently generate…

计算与语言 · 计算机科学 2026-01-21 George Mihaila , Suleyman Olcay Polat , Poli Nemkova , Himanshu Sharma , Namratha V. Urs , Mark V. Albert

Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage. Processing entire videos with…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Ruyi Xu , Guangxuan Xiao , Yukang Chen , Liuning He , Kelly Peng , Yao Lu , Song Han

The remarkable zero-shot reasoning capabilities of large-scale Visual Language Models (VLMs) on static images have yet to be fully translated to the video domain. Conventional video understanding models often rely on extensive,…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Shihao Ji , Zihui Song

The remarkable progress in text-to-video diffusion models enables the generation of photorealistic videos, although the content of these generated videos often includes unnatural movement or deformation, reverse playback, and motionless…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Yuta Oshima , Masahiro Suzuki , Yutaka Matsuo , Hiroki Furuta

The amplification of high-speed micro-motions holds significant promise, with applications spanning fault detection in fast-paced industrial environments to refining precision in medical procedures. However, conventional motion…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Baoyue Zhang , Yajing Zheng , Shiyan Chen , Jiyuan Zhang , Kang Chen , Zhaofei Yu , Tiejun Huang

Personalized video highlight detection aims to shorten a long video to interesting moments according to a user's preference, which has recently raised the community's attention. Current methods regard the user's history as holistic…

计算机视觉与模式识别 · 计算机科学 2021-09-07 Runnan Chen , Penghao Zhou , Wenzhe Wang , Nenglun Chen , Pai Peng , Xing Sun , Wenping Wang

In this work, we tackle action-scene hallucination in Video Large Language Models (Video-LLMs), where models incorrectly predict actions based on the scene context or scenes based on observed actions. We observe that existing Video-LLMs…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Kyungho Bae , Jinhyung Kim , Sihaeng Lee , Soonyoung Lee , Gunhee Lee , Jinwoo Choi

The visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as GPT-3. In particular, averaging over LLM-generated class…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Karsten Roth , Jae Myung Kim , A. Sophia Koepke , Oriol Vinyals , Cordelia Schmid , Zeynep Akata