English
Related papers

Related papers: I Can Hear You: Selective Robust Training for Deep…

200 papers

Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviate this issue, with…

Recent progress in audio generation has made it increasingly easy to create highly realistic environmental soundscapes, which can be misused to produce deceptive content, such as fake alarms, gunshots, and crowd sounds, raising concerns for…

Sound · Computer Science 2026-03-10 Han Yin , Yang Xiao , Rohan Kumar Das , Jisheng Bai , Ting Dang

Pioneering advancements in artificial intelligence, especially in genAI, have enabled significant possibilities for content creation, but also led to widespread misinformation and false content. The growing sophistication and realism of…

Artificial Intelligence · Computer Science 2024-11-14 Dinesh Srivasthav P , Badri Narayan Subudhi

The widespread use of generative AI has shown remarkable success in producing highly realistic deepfakes, posing a serious threat to various voice biometric applications, including speaker verification, voice biometrics, audio conferencing,…

Sound · Computer Science 2025-09-10 Kutub Uddin , Muhammad Umar Farooq , Awais Khan , Khalid Mahmood Malik

Deepfakes offer great potential for innovation and creativity, but they also pose significant risks to privacy, trust, and security. With a vast Hindi-speaking population, India is particularly vulnerable to deepfake-driven misinformation…

Sound · Computer Science 2024-11-26 Sukhandeep Kaur , Mubashir Buhari , Naman Khandelwal , Priyansh Tyagi , Kiran Sharma

Despite significant advances in recent years, the existing Computer-Assisted Pronunciation Training (CAPT) methods detect pronunciation errors with a relatively low accuracy (precision of 60% at 40%-80% recall). This Ph.D. work proposes…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-15 Daniel Korzekwa

The rapid advancement of speech synthesis and voice conversion technologies has raised significant security concerns in multimedia forensics. Although current detection models demonstrate impressive performance, they struggle to maintain…

Sound · Computer Science 2025-11-26 Wangjie Li , Lin Li , Qingyang Hong

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Ivan Kukanov , Jun Wah Ng

In this research study, we propose a modern artificial intelligence (AI) approach to recognize deepfake voice, also known as generative AI cloned synthetic voice. Our proposed AI technology, called AntiDeepFake, consists of all main…

Sound · Computer Science 2024-02-19 Enkhtogtokh Togootogtokh , Christian Klasen

The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific purposes. To address this gap, we introduce a novel dataset of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 David Combei , Adriana Stan , Dan Oneata , Nicolas Müller , Horia Cucu

We study the problem of learning robust acoustic models in adverse environments, characterized by a significant mismatch between training and test conditions. This problem is of paramount importance for the deployment of speech recognition…

Sound · Computer Science 2022-06-30 Dino Oglic , Zoran Cvetkovic , Peter Sollich , Steve Renals , Bin Yu

In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significant risks in…

Sound · Computer Science 2024-02-28 Karthik Sivarama Krishnan , Koushik Sivarama Krishnan

With the rapid advancement of neural audio codecs, codec-based speech generation (CoSG) systems have become highly powerful. Unfortunately, CoSG also enables the creation of highly realistic deepfake speech, making it easier to mimic an…

Most of the existing video face super-resolution (VFSR) methods are trained and evaluated on VoxCeleb1, which is designed specifically for speaker identification and the frames in this dataset are of low quality. As a consequence, the VFSR…

Image and Video Processing · Electrical Eng. & Systems 2022-05-10 Liangbin Xie. Xintao Wang , Honglun Zhang , Chao Dong , Ying Shan

The rapid advances in text-to-speech (TTS) technologies have made audio deepfakes increasingly realistic and accessible, raising significant security and trust concerns. While existing research has largely focused on detecting…

Sound · Computer Science 2026-02-03 Alabi Ahmed , Vandana Janeja , Sanjay Purushotham

With the ever-rising quality of deep generative models, it is increasingly important to be able to discern whether the audio data at hand have been recorded or synthesized. Although the detection of fake speech signals has been studied…

Sound · Computer Science 2024-06-14 Hafsa Ouajdi , Oussama Hadder , Modan Tailleur , Mathieu Lagrange , Laurie M. Heller

The INTERSPEECH 2020 Deep Noise Suppression Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical…

Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the…

Sound · Computer Science 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

With every advancement in generative AI models, forensics is under increasing pressure. The constant emergence of new generation techniques makes it impossible to collect data for each manipulation to train a deepfake detection model. Thus,…

Artificial Intelligence · Computer Science 2026-05-20 Aritra Marik , Marcel Klemt , Anna Rohrbach

As audio deepfakes transition from research artifacts to widely available commercial tools, robust biometric authentication faces pressing security threats in high-stakes industries. This paper presents a systematic empirical evaluation of…

Sound · Computer Science 2026-01-07 Mengze Hong , Di Jiang , Zeying Xie , Weiwei Zhao , Guan Wang , Chen Jason Zhang