English

MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification

Computation and Language 2026-01-14 v3

Abstract

Misinformation on the web increasingly appears in multimodal forms, combining text, images, and OCR-rendered content in ways that amplify harm to public trust and vulnerable communities. While prior fact-checking systems often rely on unimodal signals or shallow fusion strategies, modern misinformation campaigns operate across modalities and require models that can reason over subtle cross-modal inconsistencies in a transparent and responsible manner. We introduce MultiCheck, a lightweight and interpretable framework for multimodal fact verification that jointly analyzes textual, visual, and OCR evidence. At its core, MultiCheck employs a relational fusion module based on element-wise difference and product operations, allowing for explicit cross-modal interaction modeling with minimal computational overhead. A contrastive alignment objective further helps the model distinguish between supporting and refuting evidence while maintaining a small memory and energy footprint, making it suitable for low-resource deployment. Evaluated on the Factify-2 (5-class) and Mocheg (3-class) benchmarks, MultiCheck achieves huge performance improvement and remains robust under noisy OCR and missing modality conditions. Its efficiency, transparency, and real-world robustness make it well-suited for journalists, civil society organisations, and web integrity efforts working to build a safer and more trustworthy web.

Keywords

Cite

@article{arxiv.2508.05097,
  title  = {MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification},
  author = {Aditya Kishore and Gaurav Kumar and Jasabanta Patro},
  journal= {arXiv preprint arXiv:2508.05097},
  year   = {2026}
}
R2 v1 2026-07-01T04:38:32.820Z