English
Related papers

Related papers: VideoWebArena: Evaluating Long Context Multimodal …

200 papers

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Zhang , Liqiang Jing , Jia Li , Yapeng Tian , Xinya Du , Yunhui Guo , Vibhav Gogate

The objective of action quality assessment is to score sports videos. However, most existing works focus only on video dynamic information (i.e., motion information) but ignore the specific postures that an athlete is performing in a video,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Ling-An Zeng , Fa-Ting Hong , Wei-Shi Zheng , Qi-Zhi Yu , Wei Zeng , Yao-Wei Wang , Jian-Huang Lai

Understanding domain-specific theorems often requires more than just text-based reasoning; effective communication through structured visual explanations is crucial for deeper comprehension. While large language models (LLMs) demonstrate…

Artificial Intelligence · Computer Science 2025-05-27 Max Ku , Thomas Chong , Jonathan Leung , Krish Shah , Alvin Yu , Wenhu Chen

Video production workflows offer a rich and demanding arena for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this…

Cryptography and Security · Computer Science 2026-05-28 Zongheng Cao , Yi Zheng , Rui Song , Xinyu Hu

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Kunchang Li , Yali Wang , Yinan He , Yizhuo Li , Yi Wang , Yi Liu , Zun Wang , Jilan Xu , Guo Chen , Ping Luo , Limin Wang , Yu Qiao

Training video-language models is often prohibitively expensive due to the high cost of processing long frame sequences and the limited availability of annotated long videos. We present VideoWeave, a simple yet effective approach to improve…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Zane Durante , Silky Singh , Arpandeep Khatua , Shobhit Agarwal , Reuben Tan , Yong Jae Lee , Jianfeng Gao , Ehsan Adeli , Li Fei-Fei

The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understanding are limited by the inherent disconnect between…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Keliang Li , Yansong Li , Hongze Shen , Mengdi Liu , Hong Chang , Shiguang Shan

Inspired by recent trends in vision and language learning, we explore applications of attention mechanisms for visio-lingual fusion within an application to story-based video understanding. Like other video-based QA tasks, video story…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Björn Bebensee , Byoung-Tak Zhang

Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despite the strong multimodal video understanding capabilities of recent Video-LLMs, their…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Haibin He , Maoyuan Ye , Jing Zhang , Juhua Liu , Bo Du

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA).…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zhou Yu , Dejing Xu , Jun Yu , Ting Yu , Zhou Zhao , Yueting Zhuang , Dacheng Tao

Training computer-use agents requires massive amounts of GUI interaction data, but manually annotating action trajectories at scale is prohibitively expensive. We present VideoAgentTrek, a scalable pipeline that automatically mines training…

Large language model (LLM) agents are increasingly capable of automating components of machine learning development, yet existing biomedical benchmarks mainly focus on question answering, reasoning, and tool usage, or evaluate only narrow…

Computational Engineering, Finance, and Science · Computer Science 2026-05-18 Loka Li , Duzhen Zhang , Xingbo Du , Leonard Song , Zixiao Wang , Assanali Aukenov , Noel Thomas , Shakhnazar Sailaukan , Yonghan Yang , Feilong Chen , Jiahua Dong , Kun Zhang , Bin Zhang , Le Song

While existing video benchmarks largely consider specialized downstream tasks like retrieval or question-answering (QA), contemporary multimodal AI systems must be capable of well-rounded common-sense reasoning akin to human visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Kate Sanders , Benjamin Van Durme

As large language models (LLMs) evolve into sophisticated autonomous agents capable of complex software development tasks, evaluating their real-world capabilities becomes critical. While existing benchmarks like…

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical…

Software Engineering · Computer Science 2026-04-02 Zehai He , Wenyi Hong , Zhen Yang , Ziyang Pan , Mingdao Liu , Xiaotao Gu , Jie Tang

Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reducing the number of visual tokens using visual resamplers.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Peiyuan Zhang , Kaichen Zhang , Bo Li , Guangtao Zeng , Jingkang Yang , Yuanhan Zhang , Ziyue Wang , Haoran Tan , Chunyuan Li , Ziwei Liu

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct…

Computation and Language · Computer Science 2026-04-14 Yunzhe Wang , Runhui Xu , Kexin Zheng , Tianyi Zhang , Jayavibhav Niranjan Kogundi , Soham Hans , Volkan Ustun

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Chenyou Fan , Xiaofan Zhang , Shu Zhang , Wensheng Wang , Chi Zhang , Heng Huang

We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher throughput compared to traditional narrow top-down…

Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly focus on either fine-grained perception or coarse…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Seng Nam Chen , Hao Chen , Chenglam Ho , Xinyu Mao , Jinping Wang , Yu Zhang , Chao Li