中文
相关论文

相关论文: DrVideo: Document Retrieval Based Long Video Under…

200 篇论文

Video classification problem has been studied many years. The success of Convolutional Neural Networks (CNN) in image recognition tasks gives a powerful incentive for researchers to create more advanced video classification approaches. As…

计算机视觉与模式识别 · 计算机科学 2017-06-15 Manuk Akopyan , Eshsou Khashba

Video Large Language Models (VideoLLMs) have made significant strides in video understanding but struggle with long videos due to the limitations of their backbone LLMs. Existing solutions rely on length extrapolation, which is…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Xiao Wang , Qingyi Si , Jianlong Wu , Shiyu Zhu , Li Cao , Liqiang Nie

Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descriptions. This issue is…

计算与语言 · 计算机科学 2024-10-07 Jiapeng Wang , Chengyu Wang , Kunzhe Huang , Jun Huang , Lianwen Jin

Recent advances in multimodal LLMs and systems that use tools for long-video QA point to the promise of reasoning over hour-long episodes. However, many methods still compress content into lossy summaries or rely on limited toolsets,…

人工智能 · 计算机科学 2025-12-24 Runtao Liu , Ziyi Liu , Jiaqi Tang , Yue Ma , Renjie Pi , Jipeng Zhang , Qifeng Chen

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent Sensemaking), a novel…

人工智能 · 计算机科学 2026-01-13 Zixing Lin , Jiale Wang , Gee Wah Ng , Lee Onn Mak , Chan Zhi Yang Jeriel , Jun Yang Lee , Yaohao Li

Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs). However, extending input modality of LLMs to video data remains a challenging endeavor, especially for long videos. Due to…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Jiajun Liu , Yibing Wang , Hanghang Ma , Xiaoping Wu , Xiaoqi Ma , Xiaoming Wei , Jianbin Jiao , Enhua Wu , Jie Hu

Long video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos.…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Yan Shu , Zheng Liu , Peitian Zhang , Minghao Qin , Junjie Zhou , Zhengyang Liang , Tiejun Huang , Bo Zhao

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Karttikeya Mangalam , Raiymbek Akshulakov , Jitendra Malik

In recent years, online lecture videos have become an increasingly popular resource for acquiring new knowledge. Systems capable of effectively understanding/indexing lecture videos are thus highly desirable, enabling downstream tasks like…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Kangda Wei , Zhengyu Zhou , Bingqing Wang , Jun Araki , Lukas Lange , Ruihong Huang , Zhe Feng

The video reasoning ability of multimodal large language models (MLLMs) is crucial for downstream tasks like video question answering and temporal grounding. While recent approaches have explored text-based chain-of-thought (CoT) reasoning…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Haoji Zhang , Xin Gu , Jiawen Li , Chixiang Ma , Sule Bai , Chubin Zhang , Bowen Zhang , Zhichao Zhou , Dongliang He , Yansong Tang

While multimodal large language models (MLLMs) have shown remarkable success across a wide range of tasks, long-form video understanding remains a significant challenge. In this study, we focus on video understanding by MLLMs. This task is…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Daichi Yashima , Shuhei Kurita , Yusuke Oda , Komei Sugiura

There has been impressive progress in Large Multimodal Models (LMMs). Recent works extend these models to long inputs, including multi-page documents and long videos. However, the model size and performance of these long context models are…

计算机视觉与模式识别 · 计算机科学 2025-04-25 De-An Huang , Subhashree Radhakrishnan , Zhiding Yu , Jan Kautz

Video smmarization is a crucial method to reduce the time of videos which reduces the spent time to watch/review a long video. This apporach has became more important as the amount of publisehed video is increasing everyday. A single or…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Vahid Ahmadi Kalkhorani , Qingquan Zhang , Guanqun Song , Ting Zhu

In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Models (MLLMs). However, video understanding remains a…

While Large Vision-Language Models (LVLMs) have achieved substantial progress in video understanding, their application to long video reasoning is hindered by uniform frame sampling and static textual reasoning, which are inefficient and…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Zefeng He , Xiaoye Qu , Yafu Li , Siyuan Huang , Daizong Liu , Yu Cheng

The rapid advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have ushered in remarkable progress in video understanding. However, a fundamental challenge persists: effectively processing and comprehending…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Dell Zhang , Xiangyu Chen , Jixiang Luo , Mengxi Jia , Changzhi Sun , Ruilong Ren , Jingren Liu , Hao Sun , Xuelong Li

Video-Text Retrieval (VTR) aims to search for the most relevant video related to the semantics in a given sentence, and vice versa. In general, this retrieval task is composed of four successive steps: video and textual feature…

计算机视觉与模式识别 · 计算机科学 2023-02-27 Cunjuan Zhu , Qi Jia , Wei Chen , Yanming Guo , Yu Liu

Long video summarization presents significant challenges for current multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Alkesh Patel , Melis Ozyildirim , Ying-Chang Cheng , Ganesh Nagarajan

Locating specific moments within long videos (20-120 minutes) presents a significant challenge, akin to finding a needle in a haystack. Adapting existing short video (5-30 seconds) grounding methods to this problem yields poor performance.…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Tanveer Hannan , Md Mohaiminul Islam , Thomas Seidl , Gedas Bertasius

Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by computational constraints and suboptimal frame selection. We…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Wenhui Tan , Ruihua Song , Jiaze Li , Jianzhong Ju , Zhenbo Luo