中文
相关论文

相关论文: DIVER: Dynamic Iterative Visual Evidence Reasoning…

200 篇论文

In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream claim verification process. Existing general multimodal retrieval methods are often…

信息检索 · 计算机科学 2026-05-28 Zhongtian Hua , Yi Luo , Meijia Yu , Yingjie Han

In multimodal misinformation, deception usually arises not just from pixel-level manipulations in an image, but from the semantic and contextual claim jointly expressed by the image-text pair. Yet most deepfake detectors, engineered to…

计算机视觉与模式识别 · 计算机科学 2026-02-03 A S M Sharifuzzaman Sagar , Mohammed Bennamoun , Farid Boussaid , Naeha Sharif , Lian Xu , Shaaban Sahmoud , Ali Kishk

Deepfake detection is a widely researched topic that is crucial for combating the spread of malicious content, with existing methods mainly modeling the problem as classification or spatial localization. The rapid advancements in generative…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Wenbo Xu , Wei Lu , Xiangyang Luo , Jiantao Zhou

With the rapid advancement of generative AI, synthetic content across images, videos, and audio has become increasingly realistic, amplifying the risk of misinformation. Existing detection approaches predominantly focus on binary…

机器学习 · 计算机科学 2025-07-23 Xu Yang , Qi Zhang , Shuming Jiang , Yaowen Xu , Zhaofan Zou , Hao Sun , Xuelong Li

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Mi Luo , Zihui Xue , Alex Dimakis , Kristen Grauman

Previous studies on multimodal fake news detection mainly focus on the alignment and integration of cross-modal features, as well as the application of text-image consistency. However, they overlook the semantic enhancement effects of large…

多媒体 · 计算机科学 2025-07-21 Peican Zhu , Yubo Jing , Le Cheng , Bin Chen , Xiaodong Cui , Lianwei Wu , Keke Tang

The rapid advancement of generative models has intensified the challenge of detecting and interpreting visual forgeries, necessitating robust frameworks for image forgery detection while providing reasoning as well as localization. While…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Ipsita Praharaj , Yukta Butala , Badrikanath Praharaj , Yash Butala

Multimodal manipulation detection aims to simultaneously identify forged image--text pairs and localize tampered regions, yet existing methods typically rely on memorizing isolated artifacts and struggle with imperceptible manipulation…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Jun Zhou , Bingwen Hu , Yaxiong Wang , Zhedong Zheng , Yongzhen Wang , Yuchen Zhang , Ping Liu

Self-consistency methods are the core technique for improving the reasoning reliability of multimodal large language models (MLLMs). By generating multiple reasoning results through repeated sampling and selecting the best answer via…

计算与语言 · 计算机科学 2026-02-05 Xinglong Yang , Zhilin Peng , Zhanzhan Liu , Haochen Shi , Sheng-Jun Huang

The standard paradigm for fake news detection mainly utilizes text information to model the truthfulness of news. However, the discourse of online fake news is typically subtle and it requires expert knowledge to use textual information to…

计算与语言 · 计算机科学 2023-10-13 Ye Jiang , Xiaomin Yu , Yimin Wang , Xiaoman Xu , Xingyi Song , Diana Maynard

Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of text-to-image (T2I) models. However, most existing evaluation methods rely on coarse-grained metrics or static…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Fulin Shi , Wenyi Xiao , Bin Chen , Liang Din , Leilei Gan

Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the cost of tool calls or rely on localized patch-based embeddings that are insufficient to…

计算与语言 · 计算机科学 2026-04-10 Mengdan Zhu , Senhao Cheng , Liang Zhao

Fake information poses one of the major threats for society in the 21st century. Identifying misinformation has become a key challenge due to the amount of fake news that is published daily. Yet, no approach is established that addresses…

信息检索 · 计算机科学 2021-03-30 Vishwani Gupta , Katharina Beckh , Sven Giesselbach , Dennis Wegener , Tim Wirtz

While Large Language Models (LLMs) excel at reasoning on text and Vision-Language Models (VLMs) are highly effective for visual perception, applying those models for visual instruction-based planning remains a widely open problem. In this…

In current web environment, fake news spreads rapidly across online social networks, posing serious threats to society. Existing multimodal fake news detection methods can generally be classified into knowledge-based and semantic-based…

人工智能 · 计算机科学 2025-03-12 Xinqi Su , Zitong Yu , Yawen Cui , Ajian Liu , Xun Lin , Yuhao Wang , Haochen Liang , Wenhui Li , Li Shen , Xiaochun Cao

Misinformation has become a pressing issue. Fake media, in both visual and textual forms, is widespread on the web. While various deepfake detection and text fake news detection methods have been proposed, they are only designed for…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Rui Shao , Tianxing Wu , Jianlong Wu , Liqiang Nie , Ziwei Liu

The growing realism of AI-generated images produced by recent GAN and diffusion models has intensified concerns over the reliability of visual media. Yet, despite notable progress in deepfake detection, current forensic systems degrade…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Anshul Bagaria

Deepfake videos present an increasing threat to society with potentially negative impact on criminal justice, democracy, and personal safety and privacy. Meanwhile, detecting deepfakes, at scale, remains a very challenging task that often…

计算机视觉与模式识别 · 计算机科学 2024-06-24 Mulin Tian , Mahyar Khayatkhoei , Joe Mathai , Wael AbdAlmageed

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced cross-modal understanding and reasoning by incorporating Chain-of-Thought (CoT) reasoning in the semantic space. Building upon this, recent studies…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Chengzhi Liu , Yuzhe Yang , Yue Fan , Qingyue Wei , Sheng Liu , Xin Eric Wang

The landscape of social media content has evolved significantly, extending from text to multimodal formats. This evolution presents a significant challenge in combating misinformation. Previous research has primarily focused on single…

多媒体 · 计算机科学 2024-09-04 Zhe Fu , Kanlun Wang , Wangjiaxuan Xin , Lina Zhou , Shi Chen , Yaorong Ge , Daniel Janies , Dongsong Zhang