English
Related papers

Related papers: Efficient Motion-Aware Video MLLM

200 papers

Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token compression methods for VideoLLMs reveals two…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xuyang Liu , Yiyu Wang , Junpeng Ma , Linfeng Zhang

In modern multimedia systems, efficient video processing is critical, especially in resource-constrained environments such as IoT-based camera networks, autonomous platforms, and wireless sensor multimedia systems. A key bottleneck in video…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Kakia Panagidi , Stathes Hadjieftymiadis

Multimodal Large Language Models advance multimodal representation learning by acquiring transferable semantic embeddings, thereby substantially enhancing performance across a range of vision-language tasks, including cross-modal retrieval,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Da Li , Yuxiao Luo , Keping Bi , Jiafeng Guo , Wei Yuan , Biao Yang , Yan Wang , Fan Yang , Tingting Gao , Guorui Zhou

Existing codecs are designed to eliminate intrinsic redundancies to create a compact representation for compression. However, strong external priors from Multimodal Large Language Models (MLLMs) have not been explicitly explored in video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-17 Pingping Zhang , Jinlong Li , Kecheng Chen , Meng Wang , Long Xu , Haoliang Li , Nicu Sebe , Sam Kwong , Shiqi Wang

We propose Motion-Compensated Latent Semantic Canvases (MCLSC) for visual situational awareness on resource-constrained edge devices. The core idea is to maintain persistent semantic metadata in two latent canvases - a slowly accumulating…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Igor Lodin , Sergii Filatov , Vira Filatova , Dmytro Filatov

We introduce motion graph, a novel approach to the video prediction problem, which predicts future video frames from limited past data. The motion graph transforms patches of video frames into interconnected graph nodes, to comprehensively…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yiqi Zhong , Luming Liang , Bohan Tang , Ilya Zharkov , Ulrich Neumann

With the rapid development of video Multimodal Large Language Models (MLLMs), numerous benchmarks have been proposed to assess their video understanding capability. However, due to the lack of rich events in the videos, these datasets may…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Yifan Du , Kun Zhou , Yuqi Huo , Yifan Li , Wayne Xin Zhao , Haoyu Lu , Zijia Zhao , Bingning Wang , Weipeng Chen , Ji-Rong Wen

Skeleton-based action recognition has garnered significant attention due to the utilization of concise and resilient skeletons. Nevertheless, the absence of detailed body information in skeletons restricts performance, while other…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Jinfu Liu , Chen Chen , Mengyuan Liu

Video processing solutions for motion analysis are key tasks in many computer vision applications, ranging from human activity recognition to object detection. In particular, speed estimation algorithms may be relevant in contexts such as…

Image and Video Processing · Electrical Eng. & Systems 2022-11-29 Veronica Mattioli , Davide Alinovi , Riccardo Raheli

As the most essential property in a video, motion information is critical to a robust and generalized video representation. To inject motion dynamics, recent works have adopted frame difference as the source of motion information in video…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Minghao Zhu , Xiao Lin , Ronghao Dang , Chengju Liu , Qijun Chen

Video recognition has been dominated by the end-to-end learning paradigm -- first initializing a video recognition model with weights of a pretrained image model and then conducting end-to-end training on videos. This enables the video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Ziyi Lin , Shijie Geng , Renrui Zhang , Peng Gao , Gerard de Melo , Xiaogang Wang , Jifeng Dai , Yu Qiao , Hongsheng Li

The Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Bin Lin , Yang Ye , Bin Zhu , Jiaxi Cui , Munan Ning , Peng Jin , Li Yuan

In this paper, a macroblock classification method is proposed for various video processing applications involving motions. Based on the analysis of the Motion Vector field in the compressed video, we propose to classify Macroblocks of each…

Multimedia · Computer Science 2016-11-17 Weiyao Lin , Ming-Ting Sun , Hongxiang Li , Zhenzhong Chen , Wei Li , Bing Zhou

Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite recent advances in long-context MLLMs, storing and attending…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Yanlai Yang , Zhuokai Zhao , Satya Narayan Shukla , Aashu Singh , Shlok Kumar Mishra , Lizhu Zhang , Mengye Ren

Adapting Multimodal Large Language Models (MLLMs) for hour-long videos is bottlenecked by context limits. Dense visual streams saturate token budgets and exacerbate the lost-in-the-middle phenomenon. Existing heuristics, like sparse…

Implicit neural representations (INRs) have emerged as a powerful framework for continuous image representation learning. In Functa-based approaches, each image is encoded as a latent modulation vector that conditions a shared INR, enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Julia Wolleb , Cristiana Baloescu , Alicia Durrer , Hemant D. Tagare , Xenophon Papademetris

The latent representation in learned image compression encompasses channel-wise, local spatial, and global spatial correlations, which are essential for the entropy model to capture for conditional entropy minimization. Efficiently…

Image and Video Processing · Electrical Eng. & Systems 2025-10-29 Wei Jiang , Jiayu Yang , Yongqi Zhai , Feng Gao , Ronggang Wang

Embodied action planning is a core challenge in robotics, requiring models to generate precise actions from visual observations and language instructions. While video generation world models are promising, their reliance on pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Yangcheng Yu , Xin Jin , Yu Shang , Xin Zhang , Haisheng Su , Wei Wu , Yong Li

The ability to accurately interpret complex visual information is a crucial topic of multimodal large language models (MLLMs). Recent work indicates that enhanced visual perception significantly reduces hallucinations and improves…

In this work, we introduce the first framework for Motion-aware Event Suppression, which learns to filter events triggered by IMOs and ego-motion in real time. Our model jointly segments IMOs in the current event stream while predicting…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Roberto Pellerito , Nico Messikommer , Giovanni Cioffi , Marco Cannici , Davide Scaramuzza
‹ Prev 1 8 9 10 Next ›