中文
相关论文

相关论文: ImplicitAVE: An Open-Source Dataset and Multimodal…

200 篇论文

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Shuo Cao , Nan Ma , Jiayang Li , Xiaohui Li , Lihao Shao , Kaiwen Zhu , Yu Zhou , Yuandong Pu , Jiarui Wu , Jiaquan Wang , Bo Qu , Wenhai Wang , Yu Qiao , Dajuin Yao , Yihao Liu

Toxicity detection in multimodal text-image content faces growing challenges, especially with multimodal implicit toxicity, where each modality appears benign on its own but conveys hazard when combined. Multimodal implicit toxicity appears…

多媒体 · 计算机科学 2025-05-21 Shiyao Cui , Qinglin Zhang , Xuan Ouyang , Renmiao Chen , Zhexin Zhang , Yida Lu , Hongning Wang , Han Qiu , Minlie Huang

Multimodal Large Language Models (MLLMs) demonstrate impressive problem-solving abilities across a wide range of tasks and domains. However, their capacity for face understanding has not been systematically studied. To address this gap, we…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Kartik Narayan , Vibashan VS , Vishal M. Patel

We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answering,…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Zheqi He , Yesheng Liu , Jing-shu Zheng , Xuejing Li , Jin-Ge Yao , Bowen Qin , Richeng Xuan , Xi Yang

Extraction of missing attribute values is to find values describing an attribute of interest from a free text input. Most past related work on extraction of missing attribute values work with a closed world assumption with the possible set…

计算与语言 · 计算机科学 2018-10-09 Guineng Zheng , Subhabrata Mukherjee , Xin Luna Dong , Feifei Li

Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately…

Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks include overly simple or uninformative samples, complicating…

Multimodal VAEs seek to model the joint distribution over heterogeneous data (e.g.\ vision, language), whilst also capturing a shared representation across such modalities. Prior work has typically combined information from the modalities…

机器学习 · 计算机科学 2022-12-19 Tom Joy , Yuge Shi , Philip H. S. Torr , Tom Rainforth , Sebastian M. Schmon , N. Siddharth

Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Lei Li , Yuqi Wang , Runxin Xu , Peiyi Wang , Xiachong Feng , Lingpeng Kong , Qi Liu

Visualization, a domain-specific yet widely used form of imagery, is an effective way to turn complex datasets into intuitive insights, and its value depends on whether data are faithfully represented, clearly communicated, and…

计算与语言 · 计算机科学 2026-03-03 Yupeng Xie , Zhiyang Zhang , Yifan Wu , Sirong Lu , Jiayi Zhang , Zhaoyang Yu , Jinlin Wang , Sirui Hong , Bang Liu , Chenglin Wu , Yuyu Luo

Best-of-N (BoN) Average Displacement Error (ADE)/ Final Displacement Error (FDE) is the most used metric for evaluating trajectory prediction models. Yet, the BoN does not quantify the whole generated samples, resulting in an incomplete…

计算机视觉与模式识别 · 计算机科学 2022-09-13 Abduallah Mohamed , Deyao Zhu , Warren Vu , Mohamed Elhoseiny , Christian Claudel

Visual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Guozhi Tang , Lele Xie , Lianwen Jin , Jiapeng Wang , Jingdong Chen , Zhen Xu , Qianying Wang , Yaqiang Wu , Hui Li

Scientific data visualization plays a crucial role in research by enabling the direct display of complex information and assisting researchers in identifying implicit patterns. Despite its importance, the use of Large Language Models (LLMs)…

Multimodal large language models (MLLMs) demonstrate remarkable capabilities in handling complex multimodal tasks and are increasingly adopted in video understanding applications. However, their rapid advancement raises serious data privacy…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Qi Li , Runpeng Yu , Xinchao Wang

Recent studies have presented compelling evidence that large language models (LLMs) can equip embodied agents with the self-driven capability to interact with the world, which marks an initial step toward versatile robotics. However, these…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Sipeng Zheng , Jiazheng Liu , Yicheng Feng , Zongqing Lu

In medical data analysis, extracting deep insights from complex, multi-modal datasets is essential for improving patient care, increasing diagnostic accuracy, and optimizing healthcare operations. However, there is currently a lack of…

人工智能 · 计算机科学 2025-12-16 Zhenghao Zhu , Chuxue Cao , Sirui Han , Yuanfeng Song , Xing Chen , Caleb Chen Cao , Yike Guo

We introduce InternVL 2.5, an advanced multimodal large language model (MLLM) series that builds upon InternVL 2.0, maintaining its core model architecture while introducing significant enhancements in training and testing strategies as…

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall…

Large Language Models (LLMs) are widely used for downstream tasks such as tabular classification, where ensuring fairness in their outputs is critical for inclusivity, equal representation, and responsible AI deployment. This study…

计算与语言 · 计算机科学 2025-08-26 Garima Chhikara , Kripabandhu Ghosh , Abhijnan Chakraborty

Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Rang Li , Lei Li , Shuhuai Ren , Hao Tian , Shuhao Gu , Shicheng Li , Zihao Yue , Yudong Wang , Wenhan Ma , Zhe Yang , Jingyuan Ma , Zhifang Sui , Fuli Luo
‹ 上一页 1 8 9 10 下一页 ›