中文
相关论文

相关论文: EgMM-Corpus: A Multimodal Vision-Language Dataset …

200 篇论文

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper,…

人工智能 · 计算机科学 2025-01-03 Shudong Liu , Yiqiao Jin , Cheng Li , Derek F. Wong , Qingsong Wen , Lichao Sun , Haipeng Chen , Xing Xie , Jindong Wang

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image captioning and…

Recently, electroencephalography (EEG) signals have been actively incorporated to decode brain activity to visual or textual stimuli and achieve object recognition in multi-modal AI. Accordingly, endeavors have been focused on building…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Xu Zheng , Ling Wang , Kanghao Chen , Yuanhuiyi Lyu , Jiazhou Zhou , Lin Wang

The highly abstract nature of image aesthetics perception (IAP) poses significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma,…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Yipo Huang , Xiangfei Sheng , Zhichao Yang , Quan Yuan , Zhichao Duan , Pengfei Chen , Leida Li , Weisi Lin , Guangming Shi

Metaphors are pervasive in communication, making them crucial for natural language processing (NLP). Previous research on automatic metaphor processing predominantly relies on training data consisting of English samples, which often reflect…

计算与语言 · 计算机科学 2025-06-10 Senqi Yang , Dongyu Zhang , Jing Ren , Ziqi Xu , Xiuzhen Zhang , Yiliao Song , Hongfei Lin , Feng Xia

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQA, a manually…

E-commerce platforms are rich in multimodal data, featuring a variety of images that depict product details. However, this raises an important question: do these images always enhance product understanding, or can they sometimes introduce…

计算与语言 · 计算机科学 2025-11-14 Xinyi Ling , Hanwen Du , Zhihui Zhu , Xia Ning

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed on general scene…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Shravan Nayak , Kanishk Jain , Rabiul Awal , Siva Reddy , Sjoerd van Steenkiste , Lisa Anne Hendricks , Karolina Stańczak , Aishwarya Agrawal

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang

Large vision-language models (LVLMs) are increasingly deployed in globally distributed applications, such as tourism assistants, yet their ability to produce culturally appropriate responses remains underexplored. Existing multimodal safety…

计算与语言 · 计算机科学 2025-12-23 Haoyi Qiu , Kung-Hsiang Huang , Ruichen Zheng , Jiao Sun , Nanyun Peng

Large Vision-Language Models (LVLMs) have recently gained attention due to their distinctive performance and broad applicability. While it has been previously shown that their efficacy in usage scenarios involving non-Western contexts falls…

计算与语言 · 计算机科学 2025-02-20 Florian Schneider , Carolin Holtermann , Chris Biemann , Anne Lauscher

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Tianbin Li , Yanzhou Su , Wei Li , Bin Fu , Zhe Chen , Ziyan Huang , Guoan Wang , Chenglong Ma , Ying Chen , Ming Hu , Yanjun Li , Pengcheng Chen , Xiaowei Hu , Zhongying Deng , Yuanfeng Ji , Jin Ye , Yu Qiao , Junjun He

Recent advancements in large vision-language models (VLMs) have primarily focused on English, with limited attention given to other languages. To address this gap, we introduce MEENA (also known as PersianMMMU), the first dataset designed…

The rapid increase in multimedia data has spurred advancements in Multimodal Summarization with Multimodal Output (MSMO), which aims to produce a multimodal summary that integrates both text and relevant images. The inherent heterogeneity…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Yanghai Zhang , Ye Liu , Shiwei Wu , Kai Zhang , Xukai Liu , Qi Liu , Enhong Chen

Recent years have witnessed a significant interest in developing large multimodal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led to the introduction of multiple LMM benchmarks to evaluate…

计算机视觉与模式识别 · 计算机科学 2024-10-25 Sara Ghaboura , Ahmed Heakl , Omkar Thawakar , Ali Alharthi , Ines Riahi , Abduljalil Saif , Jorma Laaksonen , Fahad S. Khan , Salman Khan , Rao M. Anwer

We present AgMMU, a challenging real-world benchmark for evaluating and advancing vision-language models (VLMs) in the knowledge-intensive domain of agriculture. Unlike prior datasets that rely on crowdsourced prompts, AgMMU is distilled…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Aruna Gauba , Irene Pi , Yunze Man , Ziqi Pang , Vikram S. Adve , Yu-Xiong Wang

Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness…

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that…

‹ 上一页 1 2 3 10 下一页 ›