English
Related papers

Related papers: RW-Post: Auditable Evidence-Grounded Multimodal Fa…

200 papers

Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce \textbf{RW-Post}, a post-aligned \textbf{text--image benchmark} for real-world multimodal…

Multimedia · Computer Science 2026-05-13 Danni Xu , Shaojing Fan , Harry Cheng , Mohan Kankanhalli

Online misinformation is often multimodal in nature, i.e., it is caused by misleading associations between texts and accompanying images. To support the fact-checking process, researchers have been recently developing automatic multimodal…

The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Fanrui Zhang , Dian Li , Qiang Zhang , Jun Chen , Gang Liu , Junxiong Lin , Jiahong Yan , Jiawei Liu , Zheng-Jun Zha

The rise of disinformation on social media, especially through the strategic manipulation or repurposing of images, paired with provocative text, presents a complex challenge for traditional fact-checking methods. In this paper, we…

Multimedia · Computer Science 2025-04-11 Arka Ujjal Dey , Artemis Llabrés , Ernest Valveny , Dimosthenis Karatzas

Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-only to multimodal fact-checking, a prevailing assumption is that incorporating visual evidence…

Computation and Language · Computer Science 2026-05-14 Jaeyoon Jung , Yejun Yoon , Kunwoo Park

The World Wide Web has become a popular source for gathering information and news. Multimodal information, e.g., enriching text with photos, is typically used to convey the news more effectively or to attract attention. Photo content can…

Computation and Language · Computer Science 2020-10-26 Eric Müller-Budack , Jonas Theiner , Sebastian Diering , Maximilian Idahl , Ralph Ewerth

Multimodal manipulation detection aims to simultaneously identify forged image--text pairs and localize tampered regions, yet existing methods typically rely on memorizing isolated artifacts and struggle with imperceptible manipulation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Jun Zhou , Bingwen Hu , Yaxiong Wang , Zhedong Zheng , Yongzhen Wang , Yuchen Zhang , Ping Liu

The increasing proliferation of misinformation and its alarming impact have motivated both industry and academia to develop approaches for misinformation detection and fact checking. Recent advances on large language models (LLMs) have…

Computation and Language · Computer Science 2024-07-22 Sahar Tahmasebi , Eric Müller-Budack , Ralph Ewerth

Misinformation on the web increasingly appears in multimodal forms, combining text, images, and OCR-rendered content in ways that amplify harm to public trust and vulnerable communities. While prior fact-checking systems often rely on…

Computation and Language · Computer Science 2026-01-14 Aditya Kishore , Gaurav Kumar , Jasabanta Patro

The World Wide Web and social media platforms have become popular sources for news and information. Typically, multimodal information, e.g., image and text is used to convey information more effectively and to attract attention. While in…

Information Retrieval · Computer Science 2021-04-29 Matthias Springstein , Eric Müller-Budack , Ralph Ewerth

The reasoning-based pose estimation (RPE) benchmark has emerged as a widely adopted evaluation standard for pose-aware multimodal large language models (MLLMs). Despite its significance, we identified critical reproducibility and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Junsu Kim , Naeun Kim , Jaeho Lee , Incheol Park , Dongyoon Han , Seungryul Baek

The rapid proliferation of online misinformation threatens the stability of digital social systems and poses significant risks to public trust, policy, and safety, necessitating reliable automated fake news detection. Existing methods often…

Information Retrieval · Computer Science 2026-03-06 Roopa Bukke , Soumya Pandey , Suraj Kumar , Soumi Chattopadhyay , Chandranath Adak

Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of text-to-image (T2I) models. However, most existing evaluation methods rely on coarse-grained metrics or static…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Fulin Shi , Wenyi Xiao , Bin Chen , Liang Din , Leilei Gan

Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Bing Wang , Ximing Li , Yanjun Wang , Changchun Li , Lin Yuanbo Wu , Buyu Wang , Shengsheng Wang

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Minheng Ni , Zhengyuan Yang , Linjie Li , Chung-Ching Lin , Kevin Lin , Wangmeng Zuo , Lijuan Wang

Multimodal misinformation, encompassing textual, visual, and cross-modal distortions, poses an increasing societal threat that is amplified by generative AI. Existing methods typically focus on a single type of distortion and struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Zehong Yan , Peng Qi , Wynne Hsu , Mong Li Lee

Misinformation is now a major problem due to its potential high risks to our core democratic and societal values and orders. Out-of-context misinformation is one of the easiest and effective ways used by adversaries to spread viral false…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Sahar Abdelnabi , Rakibul Hasan , Mario Fritz

This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents across multiple…

Computation and Language · Computer Science 2026-02-23 Jaeyoon Jung , Yejun Yoon , Kunwoo Park

Misinformation is often conveyed in multiple modalities, e.g. a miscaptioned image. Multimodal misinformation is perceived as more credible by humans, and spreads faster than its text-only counterparts. While an increasing body of research…

Computation and Language · Computer Science 2023-10-27 Mubashara Akhtar , Michael Schlichtkrull , Zhijiang Guo , Oana Cocarascu , Elena Simperl , Andreas Vlachos

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sources and verifying…

Computation and Language · Computer Science 2026-05-08 David Wan , Han Wang , Ziyang Wang , Elias Stengel-Eskin , Hyunji Lee , Mohit Bansal
‹ Prev 1 2 3 10 Next ›