中文
相关论文

相关论文: Generalizable Audio Spoofing Detection using Non-S…

200 篇论文

Several types of spoofed audio, such as mimicry, replay attacks, and deepfakes, have created societal challenges to information integrity. Recently, researchers have worked with sociolinguistics experts to label spoofed audio samples with…

Audio is one of the most used ways of human communication, but at the same time it can be easily misused to trick people. With the revolution of AI, the related technologies are now accessible to almost everyone, thus making it simple for…

In recent years, the remarkable advancements in deep neural networks have brought tremendous convenience. However, the training process of a highly effective model necessitates a substantial quantity of samples, which brings huge potential…

声音 · 计算机科学 2024-09-13 Zhisheng Zhang , Pengyang Huang

Developing a reliable anomalous sound detection (ASD) system requires robustness to noise, adaptation to domain shifts, and effective performance with limited training data. Current leading methods rely on extensive labeled data for each…

声音 · 计算机科学 2024-09-10 Phurich Saengthong , Takahiro Shinozaki

In this work, we address the challenge of generalizable audio deepfake detection (ADD) across diverse speech synthesis paradigms-including conventional text-to-speech (TTS) systems and modern diffusion or flow-matching (FM) based…

音频与语音处理 · 电气工程与系统科学 2025-11-17 Farhan Sheth , Girish , Mohd Mujtaba Akhtar , Muskaan Singh

Audio plays a crucial role in applications like speaker verification, voice-enabled smart devices, and audio conferencing. However, audio manipulations, such as deepfakes, pose significant risks by enabling the spread of misinformation. Our…

声音 · 计算机科学 2025-07-18 Kutub Uddin , Awais Khan , Muhammad Umar Farooq , Khalid Malik

Audio deepfakes are increasingly in-differentiable from organic speech, often fooling both authentication systems and human listeners. While many techniques use low-level audio features or optimization black-box model training, focusing on…

声音 · 计算机科学 2025-02-21 Kevin Warren , Daniel Olszewski , Seth Layton , Kevin Butler , Carrie Gates , Patrick Traynor

Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same dataset. However, the problem remains challenging when one tries to generalize the detector to forgeries…

计算机视觉与模式识别 · 计算机科学 2022-04-04 Liang Chen , Yong Zhang , Yibing Song , Lingqiao Liu , Jue Wang

New advancements for the detection of synthetic images are critical for fighting disinformation, as the capabilities of generative AI models continuously evolve and can lead to hyper-realistic synthetic imagery at unprecedented scale and…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Pantelis Dogoulis , Giorgos Kordopatis-Zilos , Ioannis Kompatsiaris , Symeon Papadopoulos

The increasing realism of AI-generated images has raised serious concerns about misinformation and privacy violations, highlighting the urgent need for accurate and interpretable detection methods. While existing approaches have made…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Tai-Ming Huang , Wei-Tung Lin , Kai-Lung Hua , Wen-Huang Cheng , Junichi Yamagishi , Jun-Cheng Chen

Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we establish the largest public voice dataset to date, named…

Although existing face anti-spoofing (FAS) methods achieve high accuracy in intra-domain experiments, their effects drop severely in cross-domain scenarios because of poor generalization. Recently, multifarious techniques have been…

计算机视觉与模式识别 · 计算机科学 2022-01-03 Shice Liu , Shitao Lu , Hongyi Xu , Jing Yang , Shouhong Ding , Lizhuang Ma

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

声音 · 计算机科学 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

Self-supervised learning (SSL) speech models, which can serve as powerful upstream models to extract meaningful speech representations, have achieved unprecedented success in speech representation learning. However, their effectiveness on…

声音 · 计算机科学 2023-02-01 Tung-Yu Wu , Chen-An Li , Tzu-Han Lin , Tsu-Yuan Hsu , Hung-Yi Lee

As audio deepfakes transition from research artifacts to widely available commercial tools, robust biometric authentication faces pressing security threats in high-stakes industries. This paper presents a systematic empirical evaluation of…

声音 · 计算机科学 2026-01-07 Mengze Hong , Di Jiang , Zeying Xie , Weiwei Zhao , Guan Wang , Chen Jason Zhang

Spoof detectors are classifiers that are trained to distinguish spoof fingerprints from bonafide ones. However, state of the art spoof detectors do not generalize well on unseen spoof materials. This study proposes a style transfer based…

计算机视觉与模式识别 · 计算机科学 2019-12-10 Rohit Gajawada , Additya Popli , Tarang Chugh , Anoop Namboodiri , Anil K. Jain

The state-of-art models for speech synthesis and voice conversion are capable of generating synthetic speech that is perceptually indistinguishable from bonafide human speech. These methods represent a threat to the automatic speaker…

机器学习 · 计算机科学 2019-07-11 Moustafa Alzantot , Ziqi Wang , Mani B. Srivastava

Current audio deepfake detectors cannot be trusted. While they excel on controlled benchmarks, they fail when tested in the real world. We introduce Perturbed Public Voices (P$^{2}$V), an IRB-approved dataset capturing three critical…

声音 · 计算机科学 2025-08-18 Chongyang Gao , Marco Postiglione , Isabel Gortner , Sarit Kraus , V. S. Subrahmanian

Although widely adopted, existing approaches for fine-tuning pre-trained language models have been shown to be unstable across hyper-parameter settings, motivating recent work on trust region methods. In this paper, we present a simplified…

机器学习 · 计算机科学 2020-08-10 Armen Aghajanyan , Akshat Shrivastava , Anchit Gupta , Naman Goyal , Luke Zettlemoyer , Sonal Gupta

We examine the text-free speech representations of raw audio obtained from a self-supervised learning (SSL) model by analyzing the synthesized speech using the SSL representations instead of conventional text representations. Since raw…

计算与语言 · 计算机科学 2024-12-05 Joonyong Park , Daisuke Saito , Nobuaki Minematsu