中文
相关论文

相关论文: IndieFake Dataset: A Benchmark Dataset for Audio D…

200 篇论文

Thanks to recent advances in deep learning, sophisticated generation tools exist, nowadays, that produce extremely realistic synthetic speech. However, malicious uses of such tools are possible and likely, posing a serious threat to our…

声音 · 计算机科学 2022-09-29 Alessandro Pianese , Davide Cozzolino , Giovanni Poggi , Luisa Verdoliva

Deep learning (DL) has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While AudioSet is a pivotal step to bridge this gap as a…

Neural speech synthesis techniques have enabled highly realistic speech deepfakes, posing major security risks. Speech deepfake detection is challenging due to distribution shifts across spoofing methods and variability in speakers,…

声音 · 计算机科学 2025-09-30 Pu Huang , Shouguang Wang , Siya Yao , Mengchu Zhou

The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Zhixi Cai , Shreya Ghosh , Aman Pankaj Adatia , Munawar Hayat , Abhinav Dhall , Tom Gedeon , Kalin Stefanov

As increasing development of text-to-speech (TTS) and voice conversion (VC) technologies, the detection of synthetic speech has been suffered dramatically. In order to promote the development of synthetic speech detection model against…

声音 · 计算机科学 2021-10-19 Zhenyu Zhang , Yewei Gu , Xiaowei Yi , Xianfeng Zhao

Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of synthetic speech generated from these methods to commit…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Amit Kumar Singh Yadav , Kratika Bhagtani , Davide Salvi , Paolo Bestagini , Edward J. Delp

We introduce a new dataset of conversational speech representing English from India, Nigeria, and the United States. The Multi-Dialect Dataset of Dialogues (MD3) strikes a new balance between open-ended conversational speech and…

计算与语言 · 计算机科学 2023-05-22 Jacob Eisenstein , Vinodkumar Prabhakaran , Clara Rivera , Dorottya Demszky , Devyani Sharma

Facial recognition technology has made significant advances, yet its effectiveness across diverse ethnic backgrounds, particularly in specific Indian demographics, is less explored. This paper presents a detailed evaluation of both…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Pranav Pant , Niharika Dadu , Harsh V. Singh , Anshul Thakur

Speaker verification systems have been used in many production scenarios in recent years. Unfortunately, they are still highly prone to different kinds of spoofing attacks such as voice conversion and speech synthesis, etc. In this paper,…

音频与语音处理 · 电气工程与系统科学 2021-09-07 Junxiao Xue , Hao Zhou , Yabo Wang

This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs…

声音 · 计算机科学 2025-07-25 Xuechen Liu , Wanying Ge , Xin Wang , Junichi Yamagishi

Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech…

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

声音 · 计算机科学 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

Audio deepfakes pose significant threats, including impersonation, fraud, and reputation damage. To address these risks, audio deepfake detection (ADD) techniques have been developed, demonstrating success on benchmarks like ASVspoof2019.…

声音 · 计算机科学 2025-01-22 Muhammad Umar Farooq , Awais Khan , Kutub Uddin , Khalid Mahmood Malik

ASVspoof 2021 is the forth edition in the series of bi-annual challenges which aim to promote the study of spoofing and the design of countermeasures to protect automatic speaker verification systems from manipulation. In addition to a…

With the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Run Wang , Felix Juefei-Xu , Yihao Huang , Qing Guo , Xiaofei Xie , Lei Ma , Yang Liu

Automatic speaker verification (ASV) technology is recently finding its way to end-user applications for secure access to personal data, smart services or physical facilities. Similar to other biometric technologies, speaker verification is…

声音 · 计算机科学 2016-09-16 Cemal Hanilci , Tomi Kinnunen , Md Sahidullah , Aleksandr Sizov

The increasing realism of synthetic speech generated by advanced text-to-speech (TTS) models, coupled with post-processing and laundering techniques, presents a significant challenge for audio forensic detection. In this paper, we introduce…

Partial audio deepfake localization poses unique challenges and remain underexplored compared to full-utterance spoofing detection. While recent methods report strong in-domain performance, their real-world utility remains unclear. In this…

声音 · 计算机科学 2025-09-01 Hieu-Thi Luong , Inbal Rimon , Haim Permuter , Kong Aik Lee , Eng Siong Chng

Deepfakes, synthetic media created using advanced AI techniques, pose a growing threat to information integrity, particularly in politically sensitive contexts. This challenge is amplified by the increasing realism of modern generative…

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

声音 · 计算机科学 2022-06-22 Yuan Gong , Jin Yu , James Glass