中文
相关论文

相关论文: LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sa…

200 篇论文

Video large language models (Video-LLMs) have made significant progress in understanding videos. However, processing multiple frames leads to lengthy visual token sequences, presenting challenges such as the limited context length cannot…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Hui Sun , Shiyin Lu , Huanyu Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Kaifu Zhang , Ming Li

Understanding long-form videos remains a significant challenge for vision--language models (VLMs) due to their extensive temporal length and high information density. Most current multimodal large language models (MLLMs) rely on uniform…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Xian Zhang , Zexi Wu , Zinuo Li , Hongming Xu , Luqi Gong , Farid Boussaid , Naoufel Werghi , Mohammed Bennamoun

The application of Large Multimodal Models (LMMs) to long-form video understanding is constrained by limited context lengths and the computationally prohibitive cost of processing dense video tokens. Consequently, recent research has…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Jialuo Li , Bin Li , Jiahao Li , Yan Lu

Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Kai Hu , Feng Gao , Xiaohan Nie , Peng Zhou , Son Tran , Tal Neiman , Lingyun Wang , Mubarak Shah , Raffay Hamid , Bing Yin , Trishul Chilimbi

Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video comprehension due to the large volume of data and temporal…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Shaojie Zhang , Jiahui Yang , Jianqin Yin , Zhenbo Luo , Jian Luan

Understanding long-form video content presents significant challenges due to its temporal complexity and the substantial computational resources required. In this work, we propose an agent-based approach to enhance both the efficiency and…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Sullam Jeoung , Goeric Huybrechts , Bhavana Ganesh , Aram Galstyan , Sravan Bodapati

Informative data selection is a key requirement for large language models (LLMs) to minimize the amount of data required for fine-tuning, network distillation, and token pruning, enabling fast and efficient deployment, especially under…

机器学习 · 计算机科学 2026-02-03 Ahmad Sarlak , Abolfazl Razi

Recent advancements in video large language models (Video LLMs) have significantly advanced the field of video question answering (VideoQA). While existing methods perform well on short videos, they often struggle with long-range reasoning…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Mustafa Chasmai , Gauri Jagatap , Gouthaman KV , Grant Van Horn , Subhransu Maji , Andrea Fanelli

Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video segments during inference. However, two structural gaps…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Peng Zhang , Guanghao Zhang , Wanggui He , Longxiang Zhang , Mushui Liu , Yan Xia , Zhenhao Peng , Weilong Dai , Jinlong Liu , Haobing Tang , Le Zhang , Hao Jiang , Pipei Huang

Recovering a dynamic 3D scene from a long monocular video is crucial for dense geometry, camera motion, and temporal correspondence to remain consistent in a shared coordinate system. Existing methods face two key challenges: (1)…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Chenyi Xu , Yihao Wu , Liqi Yan , Chao Yang , Jianhui Zhang , Fangli Guan , Pan Li

The rise of Large Vision-Language Models (LVLMs) has significantly advanced video understanding. However, efficiently processing long videos remains a challenge due to the ``Sampling Dilemma'': low-density sampling risks missing critical…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Tianyuan Qu , Longxiang Tang , Bohao Peng , Senqiao Yang , Bei Yu , Jiaya Jia

Large video-language models (VLMs) have demonstrated promising progress in various video understanding tasks. However, their effectiveness in long-form video analysis is constrained by limited context windows. Traditional approaches, such…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Shuming Liu , Chen Zhao , Tianqi Xu , Bernard Ghanem

Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Hong Gao , Yiming Bao , Xuezhen Tu , Bin Zhong , Linan Yue , Minling Zhang

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Ziyu Ma , Chenhui Gou , Hengcan Shi , Bin Sun , Shutao Li , Hamid Rezatofighi , Jianfei Cai

Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on tasks involving images or short videos. This has led to great interest in improving the long…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Lars Doorenbos , Federico Spurio , Juergen Gall

Long-form video understanding is essential for various applications such as video retrieval, summarizing, and question answering. Yet, traditional approaches demand substantial computing power and are often bottlenecked by GPU memory. To…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Saket Gurukar , Asim Kadav

Video depth estimation lifts monocular video clips to 3D by inferring dense depth at every frame. Recent advances in single-image depth estimation, brought about by the rise of large foundation models and the use of synthetic training data,…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Bingxin Ke , Dominik Narnhofer , Shengyu Huang , Lei Ke , Torben Peters , Katerina Fragkiadaki , Anton Obukhov , Konrad Schindler

Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints and the need to capture information distributed across thousands of frames. Existing…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Junbo Zou , Ziheng Huang , Shengjie Zhang , Liwen Zhang , Weining Shen

The exponential increase in video content poses significant challenges in terms of efficient navigation, search, and retrieval, thus requiring advanced video summarization techniques. Existing video summarization methods, which heavily rely…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Min Jung Lee , Dayoung Gong , Minsu Cho

In some practical learning tasks, such as traffic video analysis, the number of available training samples is restricted by different factors, such as limited communication bandwidth and computation power. Determinantal Point Process (DPP)…

机器学习 · 计算机科学 2023-08-17 Xiwen Chen , Huayu Li , Rahul Amin , Abolfazl Razi
‹ 上一页 1 2 3 10 下一页 ›