English
Related papers

Related papers: From Talking to Singing: A New Challenge for Audio…

200 papers

Artefacts that serve to distinguish bona fide speech from spoofed or deepfake speech are known to reside in specific subbands and temporal segments. Various approaches can be used to capture and model such artefacts, however, none works…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-24 Hemlata Tak , Jee-weon Jung , Jose Patino , Madhu Kamble , Massimiliano Todisco , Nicholas Evans

Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling…

Sound · Computer Science 2025-12-09 Alireza Mohammadi , Keshav Sood , Dhananjay Thiruvady , Asef Nazari

Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on generator-specific forgery patterns, self-supervised detectors offer stronger…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Ke Liu , Jiwei Wei , Shuchang Zhou , Yutong Xiao , Ruikun Chai , Yitong Qin , Yuyang Zhou , Yang Yang

With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake…

Sound · Computer Science 2025-06-04 Chi Ding , Junxiao Xue , Cong Wang , Hao Zhou

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental sounds have…

Sound · Computer Science 2025-09-30 Han Yin , Yang Xiao , Rohan Kumar Das , Jisheng Bai , Haohe Liu , Wenwu Wang , Mark D Plumbley

Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing…

Sound · Computer Science 2025-08-11 Wei Chen , Binzhu Sha , Dan Luo , Jing Yang , Zhuo Wang , Fan Fan , Zhiyong Wu

Better generative models and larger datasets have led to more realistic fake videos that can fool the human eye but produce temporal and spatial artifacts that deep learning approaches can detect. Most current Deepfake detection methods…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Oscar de Lima , Sean Franklin , Shreshtha Basu , Blake Karwoski , Annet George

The proliferation of highly realistic singing voice deepfakes presents a significant challenge to protecting artist likeness and content authenticity. Automatic singer identification in vocal deepfakes is a promising avenue for artists and…

Sound · Computer Science 2025-11-19 Davide Salvi , Hendrik Vincent Koops , Elio Quinton

Conventional spoofing detection systems have heavily relied on the use of handcrafted features derived from speech data. However, a notable shift has recently emerged towards the direct utilization of raw speech waveforms, as demonstrated…

Entertainment-oriented singing voice synthesis (SVS) requires a vocoder to generate high-fidelity (e.g. 48kHz) audio. However, most text-to-speech (TTS) vocoders cannot reconstruct the waveform well in this scenario. In this paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Chunhui Wang , Chang Zeng , Jun Chen , Xing He

For recognizing speakers in video streams, significant research studies have been made to obtain a rich machine learning model by extracting high-level speaker's features such as facial expression, emotion, and gender. However, generating…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Ehsan Asali , Farzan Shenavarmasouleh , Farid Ghareh Mohammadi , Prasanth Sengadu Suresh , Hamid R. Arabnia

Deepfake (DF) attacks pose a growing threat as generative models become increasingly advanced. However, our study reveals that existing DF datasets fail to deceive human perception, unlike real DF attacks that influence public discourse. It…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Oguzhan Baser , Ahmet Ege Tanriverdi , Sriram Vishwanath , Sandeep P. Chinchali

A major challenge in DeepFake forgery detection is that state-of-the-art algorithms are mostly trained to detect a specific fake method. As a result, these approaches show poor generalization across different types of facial manipulations,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Davide Cozzolino , Andreas Rössler , Justus Thies , Matthias Nießner , Luisa Verdoliva

Spoofing attacks posed by generating artificial speech can severely degrade the performance of a speaker verification system. Recently, many anti-spoofing countermeasures have been proposed for detecting varying types of attacks from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-08 Yuanjun Zhao , Roberto Togneri , Victor Sreeram

Thanks to recent advances in deep learning, sophisticated generation tools exist, nowadays, that produce extremely realistic synthetic speech. However, malicious uses of such tools are possible and likely, posing a serious threat to our…

Sound · Computer Science 2022-09-29 Alessandro Pianese , Davide Cozzolino , Giovanni Poggi , Luisa Verdoliva

Existing face forgery detection usually follows the paradigm of training models in a single domain, which leads to limited generalization capacity when unseen scenarios and unknown attacks occur. In this paper, we elaborately investigate…

Computer Vision and Pattern Recognition · Computer Science 2024-07-01 Yingxin Lai , Zitong Yu , Jing Yang , Bin Li , Xiangui Kang , Linlin Shen

Voice disorders negatively impact the quality of daily life in various ways. However, accurately recognizing the category of pathological features from raw audio remains a considerable challenge due to the limited dataset. A promising…

Sound · Computer Science 2024-10-08 Lipeng Shen , Yifan Xiong , Dongyue Guo , Wei Mo , Lingyu Yu , Hui Yang , Yi Lin

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

Sound · Computer Science 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

The rapid evolution of deep generative models poses a critical challenge to deepfake detection, as detectors trained on forgery-specific artifacts often suffer significant performance degradation when encountering unseen forgeries. While…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Mengyu Qiao , Runze Tian , Yang Wang

Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain…

Sound · Computer Science 2024-07-29 Yi Zhu , Surya Koppisetti , Trang Tran , Gaurav Bharaj