English
Related papers

Related papers: From Content to Audience: A Multimodal Annotation …

200 papers

Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio,…

Multimedia · Computer Science 2026-05-15 Jianghan Chao , Jianzhang Gao , Wenhui Tan , Yuchong Sun , Ruihua Song , Liyun Ru

Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instruction data that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yunheng Li , Hengrui Zhang , Meng-Hao Guo , Wenzhao Gao , Shaoyong Jia , Shaohui Jiao , Qibin Hou , Ming-Ming Cheng

Recent Multimodal Large Language Models (MLLMs) achieve promising performance on visual and audio benchmarks independently. However, the ability of these models to process cross-modal information synchronously remains largely unexplored. We…

Artificial Intelligence · Computer Science 2026-03-12 Ziwei Zhou , Rui Wang , Zuxuan Wu , Yu-Gang Jiang

Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Sankalp Sinha , Mohammad Sadil Khan , Muhammad Usama , Shino Sam , Didier Stricker , Sk Aziz Ali , Muhammad Zeshan Afzal

The growing volume of video-based news content has heightened the need for transparent and reliable methods to extract on-screen information. Yet the variability of graphical layouts, typographic conventions, and platform-specific design…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Andrea Filiberto Lucas , Dylan Seychell

Enhancing fuel efficiency in public transportation requires the integration of complex multimodal data into interpretable, decision-relevant insights. However, traditional analytics and visualization methods often yield fragmented outputs…

Artificial Intelligence · Computer Science 2025-11-18 Zhipeng Ma , Ali Rida Bahja , Andreas Burgdorf , André Pomp , Tobias Meisen , Bo Nørregaard Jørgensen , Zheng Grace Ma

Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multimodal Large Language Models (MLLMs) have shown promising capabilities, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Yandi Wang , Libin Zhan , Ziwei Huang , Tiancheng Luo , Yuxuan Jiang , Wang Dong , Leilei Gan , Jun Chen

Pretrained vision language models (VLMs) present an opportunity to caption unlabeled 3D objects at scale. The leading approach to summarize VLM descriptions from different views of an object (Luo et al., 2023) relies on a language model…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Rishabh Kabra , Loic Matthey , Alexander Lerchner , Niloy J. Mitra

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a…

With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap,…

We propose a dedicated multimodal Judge Model designed to provide reliable, explainable evaluation across a diverse suite of tasks. Our benchmark spans text, audio, image, and video modalities, drawing from carefully sampled public datasets…

Machine Learning · Computer Science 2026-01-13 Min-Han Shih , Yu-Hsin Wu , Yu-Wei Chen

Recently, video captioning has been attracting an increasing amount of interest, due to its potential for improving accessibility and information retrieval. While existing methods rely on different kinds of visual features and model…

Computer Vision and Pattern Recognition · Computer Science 2016-12-02 Xiang Long , Chuang Gan , Gerard de Melo

We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attention-based feedforward structure in such a way that allows the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Wim Boes , Hugo Van hamme

Large multimodal models (LMMs) with advanced video analysis capabilities have recently garnered significant attention. However, most evaluations rely on traditional methods like multiple-choice questions in benchmarks such as VideoMME and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Ziyang Luo , Haoning Wu , Dongxu Li , Jing Ma , Mohan Kankanhalli , Junnan Li

Vision Language Models (VLMs) are poised to revolutionize the digital transformation of pharmacyceutical industry by enabling intelligent, scalable, and automated multi-modality content processing. Traditional manual annotation of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Suyash Mishra , Qiang Li , Srikanth Patil , Anubhav Girdhar

The rapid expansion of multimedia content has made accurately retrieving relevant videos from large collections increasingly challenging. Recent advancements in text-video retrieval have focused on cross-modal interactions, large-scale…

Computation and Language · Computer Science 2024-10-17 Donghoon Han , Eunhwan Park , Gisang Lee , Adam Lee , Nojun Kwak

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuansen Liu , Haiming Tang , Jinlong Peng , Jiangning Zhang , Xiaozhong Ji , Qingdong He , Wenbin Wu , Donghao Luo , Zhenye Gan , Junwei Zhu , Yunhang Shen , Chaoyou Fu , Chengjie Wang , Xiaobin Hu , Shuicheng Yan

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Keda Tao , Yuhua Zheng , Jia Xu , Wenjie Du , Kele Shao , Hesong Wang , Xueyi Chen , Xin Jin , Junhan Zhu , Bohan Yu , Weiqiang Wang , Jian Liu , Can Qin , Yulun Zhang , Ming-Hsuan Yang , Huan Wang

This work presents AdditiveLLM2 a multi-modal, domain adapted large language model built upon the instruction tuned variant of the Gemma 3 model using a relatively small dataset of around 50 million tokens. The dataset (AdditiveLLM2-OA)…

Machine Learning · Computer Science 2026-03-24 Peter Pak , Amir Barati Farimani

As the volume of video content online grows exponentially, the demand for moderation of unsafe videos has surpassed human capabilities, posing both operational and mental health challenges. While recent studies demonstrated the merits of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Adi Levi , Or Levi , Sardhendu Mishra , Jonathan Morra