English
Related papers

Related papers: MSA at ImageCLEF 2025 Multimodal Reasoning: Multil…

200 papers

With the rapid advancement of Artificial Intelligence (AI), Large Language Models (LLMs) have significantly impacted a wide array of domains, including healthcare, engineering, science, education, and mathematical reasoning. Among these,…

Machine Learning · Computer Science 2025-05-20 Afrar Jahin , Arif Hassan Zidan , Wei Zhang , Yu Bao , Tianming Liu

Assessing soft skills such as empathy, ethical judgment, and communication is essential in competitive selection processes, yet human scoring is often inconsistent and biased. While Large Language Models (LLMs) have improved Automated Essay…

Computation and Language · Computer Science 2026-02-03 Ryan Huynh , Frank Guerin , Alison Callwood

Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language Models (MLLMs) have demonstrated impressive reasoning…

Computation and Language · Computer Science 2026-05-19 Jen-tse Huang , Chang Chen , Shiyang Lai , Wenxuan Wang , Michelle R. Kaufman , Mark Dredze

This paper presents VisLingInstruct, a novel approach to advancing Multi-Modal Language Models (MMLMs) in zero-shot learning. Current MMLMs show impressive zero-shot abilities in multi-modal tasks, but their performance depends heavily on…

Artificial Intelligence · Computer Science 2024-06-21 Dongsheng Zhu , Xunzhu Tang , Weidong Han , Jinghui Lu , Yukun Zhao , Guoliang Xing , Junfeng Wang , Dawei Yin

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit strong reasoning…

Computation and Language · Computer Science 2025-04-25 Liangyu Xu , Yingxiu Zhao , Jingyun Wang , Yingyao Wang , Bu Pi , Chen Wang , Mingliang Zhang , Jihao Gu , Xiang Li , Xiaoyong Zhu , Jun Song , Bo Zheng

Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Linhao Yu , Xinguang Ji , Yahui Liu , Fanheng Kong , Chenxi Sun , Jingyuan Zhang , Hongzhi Zhang , V. W. , Fuzheng Zhang , Deyi Xiong

End-to-end autonomous driving systems map sensor data directly to control commands, but remain opaque, lack interpretability, and offer no formal safety guarantees. While recent vision-language-guided reinforcement learning (RL) methods…

This paper presents MSLEF, a multi-segment ensemble framework that employs LLM fine-tuning to enhance resume parsing in recruitment automation. It integrates fine-tuned Large Language Models (LLMs) using weighted voting, with each model…

Computation and Language · Computer Science 2025-09-09 Omar Walid , Mohamed T. Younes , Khaled Shaban , Mai Hassan , Ali Hamdi

Existing evaluation frameworks for Multimodal Large Language Models (MLLMs) primarily focus on image reasoning or general video understanding tasks, largely overlooking the significant role of image context in video comprehension. To bridge…

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image…

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

Computation and Language · Computer Science 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu

To advance argumentative stance prediction as a multimodal problem, the First Shared Task in Multimodal Argument Mining hosted stance prediction in crucial social topics of gun control and abortion. Our exploratory study attempts to…

Computation and Language · Computer Science 2023-10-12 Arushi Sharma , Abhibha Gupta , Maneesh Bilalpur

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Nilay Yilmaz , Maitreya Patel , Yiran Lawrence Luo , Tejas Gokhale , Chitta Baral , Suren Jayasuriya , Yezhou Yang

This report presents a solution for the zero-shot referring expression comprehension task. Visual-language multimodal base models (such as CLIP, SAM) have gained significant attention in recent years as a cornerstone of mainstream research.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Longfei Huang , Feng Yu , Zhihao Guan , Zhonghua Wan , Yang Yang

SemEval-2025 Task 1 focuses on ranking images based on their alignment with a given nominal compound that may carry idiomatic meaning in both English and Brazilian Portuguese. To address this challenge, this work uses generative large…

Computation and Language · Computer Science 2025-05-02 Thanet Markchom , Tong Wu , Liting Huang , Huizhi Liang

In this paper, we present our solution for the semi-supervised learning track (MER-SEMI) in MER2025. We propose a comprehensive framework, grounded in the principle that "more is better," to construct a robust Mixture of Experts (MoE)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Jun Xie , Yingjian Zhu , Feng Chen , Zhenghao Zhang , Xiaohui Fan , Hongzhu Yi , Xinming Wang , Chen Yu , Yue Bi , Zhaoran Zhao , Xiongjun Guan , Zhepeng Wang

Purpose: To evaluate the accuracy and reasoning ability of DeepSeek-R1 and three other recently released large language models (LLMs) in bilingual complex ophthalmology cases. Methods: A total of 130 multiple-choice questions (MCQs) related…

Computation and Language · Computer Science 2025-02-26 Pusheng Xu , Yue Wu , Kai Jin , Xiaolan Chen , Mingguang He , Danli Shi

Scene understanding is critical for various downstream tasks in autonomous driving, including facilitating driver-agent communication and enhancing human-centered explainability of autonomous vehicle (AV) decisions. This paper evaluates the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Mohammed Elhenawy , Shadi Jaradat , Taqwa I. Alhadidi , Huthaifa I. Ashqar , Ahmed Jaber , Andry Rakotonirainy , Mohammad Abu Tami

The emergence of multimodal large language models (MLLMs) has triggered extensive research in model evaluation. While existing evaluation studies primarily focus on unimodal (vision-only) comprehension and reasoning capabilities, they…

Multimedia · Computer Science 2025-04-24 Xiaocui Yang , Wenfang Wu , Shi Feng , Ming Wang , Daling Wang , Yang Li , Qi Sun , Yifei Zhang , Xiaoming Fu , Soujanya Poria

Multimodal Large Language Models (MLLMs) have showcased impressive skills in tasks related to visual understanding and reasoning. Yet, their widespread application faces obstacles due to the high computational demands during both the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Minjie Zhu , Yichen Zhu , Xin Liu , Ning Liu , Zhiyuan Xu , Chaomin Shen , Yaxin Peng , Zhicai Ou , Feifei Feng , Jian Tang