中文
相关论文

相关论文: DGM4+: Dataset Extension for Global Scene Inconsis…

200 篇论文

Misinformation has become a pressing issue. Fake media, in both visual and textual forms, is widespread on the web. While various deepfake detection and text fake news detection methods have been proposed, they are only designed for…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Rui Shao , Tianxing Wu , Jianlong Wu , Liqiang Nie , Ziwei Liu

Misinformation has become a pressing issue. Fake media, in both visual and textual forms, is widespread on the web. While various deepfake detection and text fake news detection methods have been proposed, they are only designed for…

计算机视觉与模式识别 · 计算机科学 2023-04-06 Rui Shao , Tianxing Wu , Ziwei Liu

We extend HAMMER, a state-of-the-art model for multimodal manipulation detection, to handle global scene inconsistencies such as foreground-background (FG-BG) mismatch. While HAMMER achieves strong performance on the DGM4 dataset, it…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Gagandeep Singh , Samudi Amarsinghe , Urawee Thani , Ki Fung Wong , Priyanka Singh , Xue Li

Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Changtao Miao , Yi Zhang , Man Luo , Weiwei Feng , Kaiyuan Zheng , Qi Chu , Tao Gong , Jianshu Li , Yunfeng Diao , Wei Zhou , Joey Tianyi Zhou , Xiaoshuai Hao

The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Jinjie Shen , Yaxiong Wang , Lechao Cheng , Nan Pu , Zhun Zhong

The proliferation of inflammatory or misleading "fake" news content has become increasingly common in recent years. Simultaneously, it has become easier than ever to use AI tools to generate photorealistic images depicting any scene…

计算机视觉与模式识别 · 计算机科学 2024-10-14 Runsheng Huang , Liam Dugan , Yue Yang , Chris Callison-Burch

We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Zhenxing Zhang , Yaxiong Wang , Lechao Cheng , Zhun Zhong , Dan Guo , Meng Wang

Training Scene Graph Generation (SGG) models with natural language captions has become increasingly popular due to the abundant, cost-effective, and open-world generalization supervision signals that natural language offers. However, such…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Zuyao Chen , Jinlin Wu , Zhen Lei , Zhaoxiang Zhang , Changwen Chen

Detecting and grounding multi-modal media manipulation (DGM^4) has become increasingly crucial due to the widespread dissemination of face forgery and text misinformation. In this paper, we present the Unified Frequency-Assisted transFormer…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Huan Liu , Zichang Tan , Qiang Chen , Yunchao Wei , Yao Zhao , Jingdong Wang

To tackle the threat of fake news, the task of detecting and grounding multi-modal media manipulation DGM4 has received increasing attention. However, most state-of-the-art methods fail to explore the fine-grained consistency within local…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Yiheng Li , Yang Yang , Zichang Tan , Huan Liu , Weihua Chen , Xu Zhou , Zhen Lei

Existing defect/anomaly generation methods often rely on few-shot learning, which overfits to specific defect categories due to the lack of large-scale paired defect editing data. This issue is aggravated by substantial variations in defect…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Yuanting Fan , Jun Liu , Bin-Bin Gao , Xiaochen Chen , Yuhuan Lin , Zhewei Dai , Jiawei Zhan , Chengjie Wang

Multimodal misinformation increasingly mixes realistic im-age edits with fluent but misleading text, producing persuasive posts that are difficult to verify. Existing systems usually rely on a single evidence source. Content-based detectors…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Gagandeep Singh , Samudi Amarasinghe , Priyanka Singh

The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs) have encoded rich world knowledge, they are not inherently…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Yiran He , Yun Cao , Bowen Yang , Zeyu Zhang

Face manipulation techniques have achieved significant advances, presenting serious challenges to security and social trust. Recent works demonstrate that leveraging multimodal models can enhance the generalization and interpretability of…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Ke Sun , Shen Chen , Taiping Yao , Ziyin Zhou , Jiayi Ji , Xiaoshuai Sun , Chia-Wen Lin , Rongrong Ji

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Zhihong Chen , Xuehai Bai , Yang Shi , Chaoyou Fu , Huanyu Zhang , Haotian Wang , Xiaoyan Sun , Zhang Zhang , Liang Wang , Yuanxing Zhang , Pengfei Wan , Yi-Fan Zhang

AI-generated content is becoming increasingly prevalent in the real world, leading to serious ethical and societal concerns. For instance, adversaries might exploit large multimodal models (LMMs) to create images that violate ethical or…

计算与语言 · 计算机科学 2025-04-14 Hongchao Fang , Yixin Liu , Jiangshu Du , Can Qin , Ran Xu , Feng Liu , Lichao Sun , Dongwon Lee , Lifu Huang , Wenpeng Yin

Fine-tuned autoregressive models for graph-to-sequence generation (G2S) often struggle with factual grounding and edit sensitivity. To tackle these issues, we propose a non-autoregressive diffusion framework that generates text by iterative…

计算与语言 · 计算机科学 2026-04-28 Aditya Hemant Shahane , Anuj Kumar Sirohi , Tanmoy Chakraborty , Prathosh A P , Sandeep Kumar

Talking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the urgent need for…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Xiaocan Chen , Qilin Yin , Jiarui Liu , Wei Lu , Xiangyang Luo , Jiantao Zhou

The misuse of advanced generative AI models has resulted in the widespread proliferation of falsified data, particularly forged human-centric audiovisual content, which poses substantial societal risks (e.g., financial fraud and social…

Existing work has observed that current text-to-image systems do not accurately reflect explicit spatial relations between objects such as 'left of' or 'below'. We hypothesize that this is because explicit spatial relations rarely appear in…

计算机视觉与模式识别 · 计算机科学 2024-03-04 Ander Salaberria , Gorka Azkune , Oier Lopez de Lacalle , Aitor Soroa , Eneko Agirre , Frank Keller
‹ 上一页 1 2 3 10 下一页 ›