English
Related papers

Related papers: RPRA-ADD: Forgery Trace Enhancement-Driven Audio D…

200 papers

When the task of locating manipulation regions in partially-fake audio (PFA) involves cross-domain datasets, the performance of deep learning models drops significantly due to the shift between the source and target domains. To address this…

Sound · Computer Science 2024-07-12 Siding Zeng , Jiangyan Yi , Jianhua Tao , Yujie Chen , Shan Liang , Yong Ren , Xiaohui Zhang

Modern deepfakes have evolved into localized and intermittent manipulations that require fine-grained temporal localization to mitigate severe digital security risks. The prohibitive cost of frame-level annotation makes weakly supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Midou Guo , Qilin Yin , Wei Lu , Rui Yang

In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significant risks in…

Sound · Computer Science 2024-02-28 Karthik Sivarama Krishnan , Koushik Sivarama Krishnan

The misuse of advanced generative AI models has resulted in the widespread proliferation of falsified data, particularly forged human-centric audiovisual content, which poses substantial societal risks (e.g., financial fraud and social…

Cryptography and Security · Computer Science 2025-10-28 Kangran Zhao , Yupeng Chen , Xiaoyu Zhang , Yize Chen , Weinan Guan , Baicheng Chen , Chengzhe Sun , Soumyya Kanti Datta , Qingshan Liu , Siwei Lyu , Baoyuan Wu

Audio fingerprinting (AFP) allows the identification of unknown audio content by extracting compact representations, termed audio fingerprints, that are designed to remain robust against common audio degradations. Neural AFP methods often…

Deepfake (DF) detectors face significant challenges when deployed in real-world environments, particularly when encountering test samples deviated from training data through either postprocessing manipulations or distribution shifts. We…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Hong-Hanh Nguyen-Le , Van-Tuan Tran , Dinh-Thuc Nguyen , Nhien-An Le-Khac

Recent deepfake detection methods have increasingly explored frequency domain representations to reveal manipulation artifacts that are difficult to detect in the spatial domain. However, most existing approaches rely primarily on spectral…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Zhen-Xin Lin , Shang-Kuan Chen

State-of-the-art methods for audio generation suffer from fingerprint artifacts and repeated inconsistencies across temporal and spectral domains. Such artifacts could be well captured by the frequency domain analysis over the spectrogram.…

Sound · Computer Science 2021-06-29 Yang Gao , Tyler Vuong , Mahsa Elyasi , Gaurav Bharaj , Rita Singh

Deep learning techniques have considerably improved speech processing in recent years. Speaker representations extracted by deep learning models are being used in a wide range of tasks such as speaker recognition and speech emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-25 Amirhossein Hajavi , Ali Etemad

As deepfake speech becomes common and hard to detect, it is vital to trace its source. Recent work on audio deepfake source tracing (ST) aims to find the origins of synthetic or manipulated speech. However, ST models must adapt to learn new…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Yang Xiao , Rohan Kumar Das

Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like…

The small amount of training data for many state-of-the-art deep learning-based Face Recognition (FR) systems causes a marked deterioration in their performance. Although a considerable amount of research has addressed this issue by…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Soroush Hashemifar , Abdolreza Marefat , Javad Hassannataj Joloudari , Hamid Hassanpour

In response to the rapid growth of Internet of Things (IoT) devices and rising security risks, Radio Frequency Fingerprint (RFF) has become key for device identification and authentication. However, various changing factors - beyond the RFF…

Signal Processing · Electrical Eng. & Systems 2025-08-19 Yezhuo Zhang , Zinan Zhou , Guangyu Li , Xuanpeng Li

Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work proposes a novel audio…

Sound · Computer Science 2025-06-04 Ajinkya Kulkarni , Sandipana Dowerah , Tanel Alumae , Mathew Magimai. -Doss

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Yidi Li , Hong Liu , Hao Tang

With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation process, with the final…

Face anti-spoofing (FAS) and face forgery detection play vital roles in securing face biometric systems from presentation attacks (PAs) and vicious digital manipulation (e.g., deepfakes). Despite promising performance upon large-scale data…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Zitong Yu , Rizhao Cai , Zhi Li , Wenhan Yang , Jingang Shi , Alex C. Kot

Recent studies have utilized visual large language models (VLMs) to answer not only "Is this face a forgery?" but also "Why is the face a forgery?" These studies introduced forgery-related attributes, such as forgery location and type, to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Tao Chen , Jingyi Zhang , Decheng Liu , Chunlei Peng

In this paper, we present Multi-scale Feature Aggregation Conformer (MFA-Conformer), an easy-to-implement, simple but effective backbone for automatic speaker verification based on the Convolution-augmented Transformer (Conformer). The…

Sound · Computer Science 2022-11-14 Yang Zhang , Zhiqiang Lv , Haibin Wu , Shanshan Zhang , Pengfei Hu , Zhiyong Wu , Hung-yi Lee , Helen Meng

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

Sound · Computer Science 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin