English
Related papers

Related papers: Towards multi-modal forgery representation learnin…

200 papers

Automatic deception detection is an important task that has gained momentum in computational linguistics due to its potential applications. In this paper, we propose a simple yet tough to beat multi-modal neural model for deception…

Computation and Language · Computer Science 2018-03-21 Gangeshwar Krishnamurthy , Navonil Majumder , Soujanya Poria , Erik Cambria

Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object localization in videos.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-11 Sizhe Li , Yapeng Tian , Chenliang Xu

Numerous synthesized videos from generative models, especially human-centric ones that simulate realistic human actions, pose significant threats to human information security and authenticity. While progress has been made in binary forgery…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chang Liu , Yunfan Ye , Fan Zhang , Qingyang Zhou , Yuchuan Luo , Zhiping Cai

With the recent advancement in large language models (LLMs), there is a growing interest in combining LLMs with multimodal learning. Previous surveys of multimodal large language models (MLLMs) mainly focus on multimodal understanding. This…

The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zhixi Cai , Shreya Ghosh , Aman Pankaj Adatia , Munawar Hayat , Abhinav Dhall , Tom Gedeon , Kalin Stefanov

The proliferation of generative AI has led to hyper-realistic synthetic videos, escalating misuse risks and outstripping binary real/fake detectors. We introduce SAGA (Source Attribution of Generative AI videos), the first comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Rohit Kundu , Vishal Mohanty , Hao Xiong , Shan Jia , Athula Balachandran , Amit K. Roy-Chowdhury

With the development of audio deepfake techniques, attacks with partially deepfake audio are beginning to rise. Compared to fully deepfake, it is much harder to be identified by the detector due to the partially cryptic manipulation,…

Sound · Computer Science 2025-07-08 Jiayi He , Jiangyan Yi , Jianhua Tao , Siding Zeng , Hao Gu

Detecting partial deepfake speech is essential due to its potential for subtle misinformation. However, existing methods depend on costly frame-level annotations during training, limiting real-world scalability. Also, they focus on…

Sound · Computer Science 2025-07-28 Menglu Li , Xiao-Ping Zhang , Lian Zhao

This paper provides a review on representation learning for videos. We classify recent spatiotemporal feature learning methods for sequential visual data and compare their pros and cons for general video analysis. Building effective…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Elham Ravanbakhsh , Yongqing Liang , J. Ramanujam , Xin Li

The rapid development of generative artificial intelligence (AI) has introduced significant opportunities for enhancing the efficiency and accuracy of image transmission within semantic communication systems. Despite these advancements,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Qiyu Ma , Wanli Ni , Zhijin Qin

Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns about digital media authenticity. While existing detection methods demonstrate promising…

Multimedia · Computer Science 2025-04-15 Junhao Xu , Jingjing Chen , Yang Jiao , Jiacheng Zhang , Zhiyu Tan , Hao Li , Yu-Gang Jiang

Recent advancements in deep learning generative models have raised concerns as they can create highly convincing counterfeit images and videos. This poses a threat to people's integrity and can lead to social instability. To address this…

Computer Vision and Pattern Recognition · Computer Science 2024-02-19 Leandro A. Passos , Danilo Jodas , Kelton A. P. da Costa , Luis A. Souza Júnior , Douglas Rodrigues , Javier Del Ser , David Camacho , João Paulo Papa

With the exponential increase in video content, the need for accurate deception detection in human-centric video analysis has become paramount. This research focuses on the extraction and combination of various features to enhance the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Mohamed Bahaa , Mena Hany , Ehab E. Zakaria

Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identity - an increasingly…

Sound · Computer Science 2026-05-20 Tung Vu , Yen Nguyen , Hai Nguyen , Cuong Pham , Cong Tran

Detecting manipulated facial images and videos is an increasingly important topic in digital media forensics. As advanced face synthesis and manipulation methods are made available, new types of fake face representations are being created…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Hao Dang , Feng Liu , Joel Stehouwer , Xiaoming Liu , Anil Jain

We present a novel algorithm for transferring artistic styles of semantically meaningful local regions of an image onto local regions of a target video while preserving its photorealism. Local regions may be selected either fully…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Xide Xia , Tianfan Xue , Wei-sheng Lai , Zheng Sun , Abby Chang , Brian Kulis , Jiawen Chen

With the advancement of deep learning-driven video editing technology, security risks have emerged. Malicious video tampering can lead to public misunderstanding, property losses, and legal disputes. Currently, detection methods are mostly…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Pengfei Pei

Human-imitated speech poses a greater challenge than AI-generated speech for both human listeners and automatic detection systems. Unlike AI-generated speech, which often contains artifacts, over-smoothed spectra, or robotic cues, imitated…

Sound · Computer Science 2026-04-28 Khalid Zaman , Masashi Unoki

The rapid proliferation of AI-powered video generation systems has introduced significant challenges in content moderation, particularly with respect to adult and sexually explicit material. Existing detection methods operate on either…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Alizishaan Khatri , Chiquita Prabhu

With the rise of generative AI technology, anyone can now easily create and deploy AI-generated music, which has heightened the need for technical solutions to address copyright and ownership issues. While existing works mainly focused on…

Sound · Computer Science 2026-01-21 Yumin Kim , Seonghyeon Go