English
Related papers

Related papers: QAMO: Quality-aware Multi-centroid One-class Learn…

200 papers

Audio DeepFakes allow the creation of high-quality, convincing utterances and therefore pose a threat due to its potential applications such as impersonation or fake news. Methods for detecting these manipulations should be characterized by…

Sound · Computer Science 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

There are growing implications surrounding generative AI in the speech domain that enable voice cloning and real-time voice conversion from one individual to another. This technology poses a significant ethical threat and could lead to…

Sound · Computer Science 2023-08-25 Jordan J. Bird , Ahmad Lotfi

Deepfake speech utterances can be forged by replacing one or more words in a bona fide utterance with semantically different words synthesized with speech-generative models. While a dedicated synthetic word detector could be developed, we…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Hoan My Tran , Xin Wang , Wanying Ge , Xuechen Liu , Junichi Yamagishi

Much of text-to-speech research relies on human evaluation, which incurs heavy costs and slows down the development process. The problem is particularly acute in heavily multilingual applications, where recruiting and polling judges can…

Computation and Language · Computer Science 2023-06-02 Thibault Sellam , Ankur Bapna , Joshua Camp , Diana Mackinnon , Ankur P. Parikh , Jason Riesa

The paper introduces Diff-Filter, a multichannel speech enhancement approach based on the diffusion probabilistic model, for improving speaker verification performance under noisy and reverberant conditions. It also presents a new two-step…

Sound · Computer Science 2023-07-06 Sandipana Dowerah , Ajinkya Kulkarni , Romain Serizel , Denis Jouvet

The proliferation of synthetic facial imagery has intensified the need for robust Open-World DeepFake Attribution (OW-DFA), which aims to attribute both known and unknown forgeries using labeled data for known types and unlabeled data…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Haiyang Zheng , Nan Pu , Wenjing Li , Teng Long , Nicu Sebe , Zhun Zhong

Despite their black-box nature, deep learning models are extensively used in image-based drug discovery to extract feature vectors from single cells in microscopy images. To better understand how these networks perform representation…

Image and Video Processing · Electrical Eng. & Systems 2024-03-27 Vivek Gopalakrishnan , Jingzhe Ma , Zhiyong Xie

This paper targets on the problem of set to set recognition, which learns the metric between two image sets. Images in each set belong to the same identity. Since images in a set can be complementary, they hopefully lead to higher accuracy…

Computer Vision and Pattern Recognition · Computer Science 2017-04-12 Yu Liu , Junjie Yan , Wanli Ouyang

Imagined speech is spotlighted as a new trend in the brain-machine interface due to its application as an intuitive communication tool. However, previous studies have shown low classification performance, therefore its use in real-life is…

Signal Processing · Electrical Eng. & Systems 2020-08-31 Dong-Yeon Lee , Minji Lee , Seong-Whan Lee

With the rise in manipulated media, deepfake detection has become an imperative task for preserving the authenticity of digital content. In this paper, we present a novel multi-modal audio-video framework designed to concurrently process…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Aaditya Kharel , Manas Paranjape , Aniket Bera

In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound scene and sound event in an input audio recording is fake or not. To this end, we conducted…

Sound · Computer Science 2026-05-04 Lam Pham , Khoi Vu , Dat Tran , Phat Lam , Vu Nguyen , David Fischinger , Son Le

The rise of advanced large language models such as GPT-4, GPT-4o, and the Claude family has made fake audio detection increasingly challenging. Traditional fine-tuning methods struggle to keep pace with the evolving landscape of synthetic…

Sound · Computer Science 2024-08-14 Xiaohui Zhang , Jiangyan Yi , Jianhua Tao

Federated learning is a paradigm that enables local devices to jointly train a server model while keeping the data decentralized and private. In federated learning, since local data are collected by clients, it is hardly guaranteed that the…

Machine Learning · Computer Science 2022-03-01 Seunghan Yang , Hyoungseob Park , Junyoung Byun , Changick Kim

It is critical for a keyword spotting model to have a small footprint as it typically runs on-device with low computational resources. However, maintaining the previous SOTA performance with reduced model size is challenging. In addition, a…

Sound · Computer Science 2022-04-13 Dianwen Ng , Jin Hui Pang , Yang Xiao , Biao Tian , Qiang Fu , Eng Siong Chng

Speech deepfake detection is a well-established research field with different models, datasets, and training strategies. However, the lack of standardized implementations and evaluation protocols limits reproducibility, benchmarking, and…

Automatic speech quality assessment is essential for audio researchers, developers, speech and language pathologists, and system quality engineers. The current state-of-the-art systems are based on framewise speech features (hand-engineered…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-15 Karl El Hajal , Zihan Wu , Neil Scheidwasser-Clow , Gasser Elbanna , Milos Cernak

Voice deepfake attacks, which artificially impersonate human speech for malicious purposes, have emerged as a severe threat. Existing defenses typically inject noise into human speech to compromise voice encoders in speech synthesis models.…

Sound · Computer Science 2025-08-26 Yuanda Wang , Bocheng Chen , Hanqing Guo , Guangjing Wang , Weikang Ding , Qiben Yan

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Hashim Ali , Surya Subramani , Lekha Bollinani , Nithin Sai Adupa , Sali El-Loh , Hafiz Malik

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively…

Sound · Computer Science 2025-06-10 Kuiyuan Zhang , Wenjie Pei , Rushi Lan , Yifang Guo , Zhongyun Hua

Current tokenization methods process sequential data without accounting for signal quality, limiting their effectiveness on noisy real-world corpora. We present QA-Token (Quality-Aware Tokenization), which incorporates data reliability…

Artificial Intelligence · Computer Science 2026-02-09 Arvid E. Gollwitzer , Paridhi Latawa , David de Gruijl , Deepak A. Subramanian , Adrián Noriega de la Colina