English
Related papers

Related papers: VTAgent: Agentic Keyframe Anchoring for Evidence-A…

200 papers

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 SunYoung Park , Jong-Hyeon Lee , Youngjune Kim , Daegyu Sung , Younghyun Yu , Young-rok Cha , Jeongho Ju

Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computation and involve unnecessary visual noise, especially in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Zheng Cheng , Rendong Wang , Zhicheng Wang

Recently, image-based Large Multimodal Models (LMMs) have made significant progress in video question-answering (VideoQA) using a frame-wise approach by leveraging large-scale pretraining in a zero-shot manner. Nevertheless, these models…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Chuyi Shang , Amos You , Sanjay Subramanian , Trevor Darrell , Roei Herzig

Videos convey rich information. Dynamic spatio-temporal relationships between people/objects, and diverse multimodal events are present in a video clip. Hence, it is important to develop automated models that can accurately extract such…

Computation and Language · Computer Science 2020-05-14 Hyounghun Kim , Zineng Tang , Mohit Bansal

An ideal vision-language agent serves as a bridge between the human users and their surrounding physical world in real-world applications like autonomous driving and embodied agents, and proactively provides accurate and timely responses…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Gengyuan Zhang , Tanveer Hannan , Hermine Kleiner , Beste Aydemir , Xinyu Xie , Jian Lan , Thomas Seidl , Volker Tresp , Jindong Gu

The rise of short-form video platforms and the emergence of multimodal large language models (MLLMs) have amplified the need for scalable, effective, zero-shot text-to-video retrieval systems. While recent advances in large-scale…

Information Retrieval · Computer Science 2026-02-24 Jiaxin Wu , Xiao-Yong Wei , Qing Li

Long video understanding is challenging due to dense visual redundancy, long-range temporal dependencies, and the tendency of chain-of-thought and retrieval-based agents to accumulate semantic drift and correlation-driven errors. We argue…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Zheng Wang , Haoran Chen , Haoxuan Qin , Zhipeng Wei , Tianwen Qian , Cong Bai

We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to focus on surface…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Aaron Foss , Chloe Evans , Sasha Mitts , Koustuv Sinha , Ammar Rizvi , Justine T. Kao

Owing to powerful natural language processing and generative capabilities, large language model (LLM) agents have emerged as a promising solution for enhancing recommendation systems via user simulation. However, in the realm of video…

Multimedia · Computer Science 2025-07-04 Siran Chen , Boyu Chen , Chenyun Yu , Yuxiao Luo , Ouyang Yi , Lei Cheng , Chengxiang Zhuo , Zang Li , Yali Wang

Significant advancements in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although these image-language models can efficiently represent both video and language branches,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Bo Zou , Chao Yang , Yu Qiao , Chengbin Quan , Youjian Zhao

Video question-answering is a fundamental task in the field of video understanding. Although current vision--language models (VLMs) equipped with Video Transformers have enabled temporal modeling and yielded superior results, they are at…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Wei Han , Hui Chen , Min-Yen Kan , Soujanya Poria

Vision-language models (VLMs) excel at extracting and reasoning about information from images. Yet, their capacity to leverage internal knowledge about specific entities remains underexplored. This work investigates the disparity in model…

Computation and Language · Computer Science 2026-01-06 Ido Cohen , Daniela Gottesman , Mor Geva , Raja Giryes

Text-based VQA aims at answering questions by reading the text present in the images. It requires a large amount of scene-text relationship understanding compared to the VQA task. Recent studies have shown that the question-answer pairs in…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Shamanthak Hegde , Soumya Jahagirdar , Shankar Gangisetty

Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video frames and their associated information. Efficiently…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Xuyi Yang , Wenhao Zhang , Hongbo Jin , Lin Liu , Hongbo Xu , Yongwei Nie , Fei Yu , Fei Ma

Visual question answering (VQA) and image captioning require a shared body of general knowledge connecting language and vision. We present a novel approach to improve VQA performance that exploits this connection by jointly generating…

Computer Vision and Pattern Recognition · Computer Science 2020-01-07 Jialin Wu , Zeyuan Hu , Raymond J. Mooney

Retrieval-Augmented Generation (RAG) is a powerful strategy for improving the factual accuracy of models by retrieving external knowledge relevant to queries and incorporating it into the generation process. However, existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Soyeong Jeong , Kangsan Kim , Jinheon Baek , Sung Ju Hwang

Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of these latent tokens at inference remains ambiguous. We show…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Dongyao Zhu , Zhen Wang , Xi Xiao , Han Jiang , Saeed Vahidian , Wei-Lun Chao , Tanya Berger-Wolf , Yu Su , Raju Vatsavai , Jianyang Gu

Visual question answering (VQA) has witnessed great progress since May, 2015 as a classic problem unifying visual and textual data into a system. Many enlightening VQA works explore deep into the image and question encodings and fusing…

Computer Vision and Pattern Recognition · Computer Science 2017-02-23 Yuetan Lin , Zhangyang Pang , Donghui Wang , Yueting Zhuang

Video Question Answering (VideoQA), aiming to correctly answer the given question based on understanding multi-modal video content, is challenging due to the rich video content. From the perspective of video understanding, a good VideoQA…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Jingjing Jiang , Ziyi Liu , Nanning Zheng

Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts. To…

Artificial Intelligence · Computer Science 2026-01-08 Rachneet Kaur , Nishan Srishankar , Zhen Zeng , Sumitra Ganesh , Manuela Veloso