English
Related papers

Related papers: Audio-centric Video Understanding Benchmark withou…

200 papers

Large Multimodal Models (LMMs) have demonstrated impressive performance in short video understanding tasks but face great challenges when applied to long video understanding. In contrast, Large Language Models (LLMs) exhibit outstanding…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Hongchen Wei , Zhenzhong Chen

Large language models (LLMs) have demonstrated remarkable capabilities in natural language and multimodal domains. By fine-tuning multimodal LLMs with temporal annotations from well-annotated datasets, e.g., dense video captioning datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Yolo Yunlong Tang , Daiki Shimada , Jing Bi , Mingqian Feng , Hang Hua , Chenliang Xu

Large multimodal models (LMMs) have shown remarkable progress in audio-visual understanding, yet they struggle with real-world scenarios that require complex reasoning across extensive video collections. Existing benchmarks for video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Sanjoy Chowdhury , Mohamed Elmoghany , Yohan Abeysinghe , Junjie Fei , Sayan Nag , Salman Khan , Mohamed Elhoseiny , Dinesh Manocha

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Yaning Pan , Qianqian Xie , Guohui Zhang , Zekun Wang , Yongqian Wen , Yuanxing Zhang , Haoxuan Hu , Zhiyu Pan , Yibing Huang , Zhidong Gan , Yonghong Lin , An Ping , Shihao Li , Yanghai Wang , Tianhao Peng , Jiaheng Liu

Recent advancements in language multimodal models (LMMs) for video have demonstrated their potential for understanding video content, yet the task of comprehending multi-discipline lectures remains largely unexplored. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Enxin Song , Wenhao Chai , Weili Xu , Jianwen Xie , Yuxuan Liu , Gaoang Wang

The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Kyungbok Lee , You Zhang , Zhiyao Duan

Understanding audio-visual content and the ability to have an informative conversation about it have both been challenging areas for intelligent systems. The Audio Visual Scene-aware Dialog (AVSD) challenge, organized as a track of the…

Computation and Language · Computer Science 2018-12-19 Dat Tien Nguyen , Shikhar Sharma , Hannes Schulz , Layla El Asri

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily generalizable to diverse tasks, domains, and datasets. To…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Linjie Li , Jie Lei , Zhe Gan , Licheng Yu , Yen-Chun Chen , Rohit Pillai , Yu Cheng , Luowei Zhou , Xin Eric Wang , William Yang Wang , Tamara Lee Berg , Mohit Bansal , Jingjing Liu , Lijuan Wang , Zicheng Liu

Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the…

Existing video understanding benchmarks often conflate knowledge-based and purely image-based questions, rather than clearly isolating a model's temporal reasoning ability, which is the key aspect that distinguishes video understanding from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Bo Feng , Zhengfeng Lai , Shiyu Li , Zizhen Wang , Simon Wang , Ping Huang , Meng Cao

Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse…

Multimedia · Computer Science 2025-03-18 Kyeongha Rho , Hyeongkeun Lee , Valentio Iverson , Joon Son Chung

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Guangyao Li , Yake Wei , Yapeng Tian , Chenliang Xu , Ji-Rong Wen , Di Hu

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant limitations in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Qi Wu , Quanlong Zheng , Yanhao Zhang , Junlin Xie , Jinguo Luo , Kuo Wang , Peng Liu , Qingsong Xie , Ru Zhen , Zhenyu Yang , Haonan Lu

Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervised approaches that use predefined labels, ALMs leverage…

Sound · Computer Science 2026-03-13 Yi Su , Jisheng Bai , Qisheng Xu , Kele Xu , Yong Dou

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often lack this capability,…

Computation and Language · Computer Science 2026-01-29 Hyunjong Ok , Suho Yoo , Hyeonjun Kim , Jaeho Lee

Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yuanhan Zhang , Yunice Chew , Yuhao Dong , Aria Leo , Bo Hu , Ziwei Liu

In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representation learning with minimal modality-specific design, and do…

Computer Vision and Pattern Recognition · Computer Science 2022-11-03 Zineng Tang , Jaemin Cho , Yixin Nie , Mohit Bansal

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent Sensemaking), a novel…

Artificial Intelligence · Computer Science 2026-01-13 Zixing Lin , Jiale Wang , Gee Wah Ng , Lee Onn Mak , Chan Zhi Yang Jeriel , Jun Yang Lee , Yaohao Li

With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yulin Fei , Yuhui Gao , Xingyuan Xian , Xiaojin Zhang , Tao Wu , Wei Chen