中文
相关论文

相关论文: Head-wise Modality Specialization within MLLMs for…

200 篇论文

This study evaluates the effectiveness of Vision Language Models (VLMs) in representing and utilizing multimodal content for fact-checking. To be more specific, we investigate whether incorporating multimodal content improves performance…

计算与语言 · 计算机科学 2024-12-09 Recep Firat Cekinel , Pinar Karagoz , Cagri Coltekin

The rapid rise of deepfake technology poses a severe threat to social and political stability by enabling hyper-realistic synthetic media capable of manipulating public perception. However, existing detection methods struggle with two core…

计算与语言 · 计算机科学 2026-01-27 Gautam Siddharth Kashyap , Harsh Joshi , Niharika Jain , Ebad Shabbir , Jiechao Gao , Nipun Joshi , Usman Naseem

Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate relevant evidence across interleaved text and images. Prior…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Aaron Branson Cigres Li , Zhaowei Wang , Yu Zhao , Yiming Du , Haobo Li , Xiyu Ren , Ginny Wong , Simon See , Lishu Luo , Haodong Duan , Pasquale Minervini , Yangqiu Song

Multimodal information processing has become increasingly important for enhancing image classification performance. However, the intricate and implicit dependencies across different modalities often hinder conventional methods from…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Yang Qiao , Xiaoyu Zhong , Xiaofeng Gu , Zhiguo Yu

With the exponential increase in video content, the need for accurate deception detection in human-centric video analysis has become paramount. This research focuses on the extraction and combination of various features to enhance the…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Mohamed Bahaa , Mena Hany , Ehab E. Zakaria

In this paper, we propose SimMLM, a simple yet powerful framework for multimodal learning with missing modalities. Unlike existing approaches that rely on sophisticated network architectures or complex data imputation techniques, SimMLM…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Sijie Li , Chen Chen , Jungong Han

In recent years, fake news detection has received increasing attention in public debate and scientific research. Despite advances in detection techniques, the production and spread of false information have become more sophisticated, driven…

计算与语言 · 计算机科学 2026-03-27 Pietro Dell'Oglio , Alessandro Bondielli , Francesco Marcelloni , Lucia C. Passaro

As medical diagnoses increasingly leverage multimodal data, machine learning models are expected to effectively fuse heterogeneous information while remaining robust to missing modalities. In this work, we propose a novel multimodal…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Yi Gu , Kuniaki Saito , Jiaxin Ma

The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Traditional unimodal detection methods fall short in addressing complex…

多媒体 · 计算机科学 2025-04-15 Moyang Liu , Kaiying Yan , Yukun Liu , Ruibo Fu , Zhengqi Wen , Xuefei Liu , Chenxing Li

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges. Existing Group Relative…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Jingyi Liao , Yongyi Su , Rong-Cheng Tu , Zhao Jin , Wenhao Sun , Yiting Li , Dacheng Tao , Xun Xu , Xulei Yang

Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial context, they lead readers to form judgments that diverge from what the full article supports.…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Fanxiao Li , Jiaying Wu , Tingchao Fu , Dayang Li , Herun Wan , Wei Zhou , Min-Yen Kan

The availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on…

计算机视觉与模式识别 · 计算机科学 2023-05-08 Ajian Liu , Zichang Tan , Zitong Yu , Chenxu Zhao , Jun Wan , Yanyan Liang , Zhen Lei , Du Zhang , Stan Z. Li , Guodong Guo

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: how can language models, initially trained solely on…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Jing Bi , Junjia Guo , Yunlong Tang , Lianggong Bruce Wen , Zhang Liu , Chenliang Xu

Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted…

计算与语言 · 计算机科学 2026-01-29 Bin Zhu , Munan Ning , Peng Jin , Bin Lin , Jinfa Huang , Qi Song , Junwu Zhang , Zhenyu Tang , Mingjun Pan , Li Yuan

In this paper, we address a fundamental gap between pre-training and fine-tuning of deep neural networks: while pre-training has shifted from unimodal to multimodal learning with enhanced visual understanding, fine-tuning predominantly…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Shohei Enomoto , Shin'ya Yamaguchi

Multimodal Large Language Models (MLLMs) extend foundation models to real-world applications by integrating inputs such as text and vision. However, their broad knowledge capacity raises growing concerns about privacy leakage, toxicity…

机器学习 · 计算机科学 2025-11-11 Kunhao Li , Wenhao Li , Di Wu , Lei Yang , Jun Bai , Ju Jia , Jason Xue

Vision and language models (VL) are known to exploit unrobust indicators in individual modalities (e.g., introduced by distributional biases) instead of focusing on relevant information in each modality. That a unimodal model achieves…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Letitia Parcalabescu , Anette Frank

We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base. Existing work lacks sufficient experimentation on the…

Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long videos. Existing methods typically rely on complex,…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Sanghwan Kim , Rui Xiao , Stephan Alaniz , Yongqin Xian , Zeynep Akata

Multi-modal data provides abundant and diverse object information, crucial for effective modal interactions in Re-Identification (ReID) tasks. However, existing approaches often overlook the quality variations in local features and fail to…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Xixi Wan , Aihua Zheng , Zi Wang , Bo Jiang , Jin Tang , Jixin Ma