English
Related papers

Related papers: HQ-MPSD: A Multilingual Artifact-Controlled Benchm…

200 papers

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Siyin Wang , Wenyi Yu , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Lu Lu , Yu Tsao , Junichi Yamagishi , Yuxuan Wang , Chao Zhang

ASVspoof 2021 is the forth edition in the series of bi-annual challenges which aim to promote the study of spoofing and the design of countermeasures to protect automatic speaker verification systems from manipulation. In addition to a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Junichi Yamagishi , Xin Wang , Massimiliano Todisco , Md Sahidullah , Jose Patino , Andreas Nautsch , Xuechen Liu , Kong Aik Lee , Tomi Kinnunen , Nicholas Evans , Héctor Delgado

Anti-spoofing is the task of speech authentication. That is, identifying genuine human speech compared to spoofed speech. The main focus of this paper is to suggest new representations for genuine and spoofed speech, based on the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-28 Matan Karo , Arie Yeredor , Itshak Lapidot

Full-duplex spoken dialogue systems (FDSDS) enable more natural human-machine interactions by allowing real-time user interruptions and backchanneling, compared to traditional SDS that rely on turn-taking. However, existing benchmarks lack…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-28 Yizhou Peng , Yi-Wen Chao , Dianwen Ng , Yukun Ma , Chongjia Ni , Bin Ma , Eng Siong Chng

Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment…

Sound · Computer Science 2023-01-27 Jianwei Zhang , Julie Liss , Suren Jayasuriya , Visar Berisha

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations, or unrealistic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Tianxiao Li , Zhenglin Huang , Haiquan Wen , Yiwei He , Xinze Li , Bingyu Zhu , Wuhui Duan , Congang Chen , Zeyu Fu , Yi Dong , Baoyuan Wu , Jason Li , Guangliang Cheng

Deepfakes have become a universal and rapidly intensifying concern of generative AI across various media types such as images, audio, and videos. Among these, audio deepfakes have been of particular concern due to the ease of high-quality…

Cryptography and Security · Computer Science 2025-03-25 Xiang Li , Pin-Yu Chen , Wenqi Wei

Deepfake detection remains highly challenging, particularly in cross-dataset scenarios and complex real-world settings. This challenge mainly arises because artifact patterns vary substantially across different forgery methods, whereas…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Xiang Zhang , Wenliang Weng , Daoyong Fu , Beijing Chen , Ziqiang Li , Ziwen He , Zhangjie Fu

AI-synthesized speech, also known as deepfake speech, has recently raised significant concerns due to the rapid advancement of speech synthesis and speech conversion techniques. Previous works often rely on distinguishing synthesizer…

Sound · Computer Science 2024-11-15 Kuiyuan Zhang , Zhongyun Hua , Yushu Zhang , Yifang Guo , Tao Xiang

Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.In this paper, we aim to fill in the gap and design a…

Sound · Computer Science 2023-07-19 Haoxin Ma , Jiangyan Yi , Chenglong Wang , Xinrui Yan , Jianhua Tao , Tao Wang , Shiming Wang , Ruibo Fu

The recent emergence of deepfakes has brought manipulated and generated content to the forefront of machine learning research. Automatic detection of deepfakes has seen many new machine learning techniques, however, human detection…

Human-Computer Interaction · Computer Science 2024-08-28 Nicolas M. Müller , Karla Pizzi , Jennifer Williams

We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to span a wide range of phonemic contrasts. Given only 10 hours…

Computation and Language · Computer Science 2026-03-20 Maxime Poli , Manel Khentout , Angelo Ortiz Tandazo , Ewan Dunbar , Emmanuel Chemla , Emmanuel Dupoux

Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific…

Sound · Computer Science 2025-10-30 Yassine El Kheir , Fabian Ritter-Guttierez , Arnab Das , Tim Polzehl , Sebastian Möller

Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest…

Deepfake generation methods are evolving fast, making fake media harder to detect and raising serious societal concerns. Most deepfake detection and dataset creation research focuses on monolingual content, often overlooking the challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kartik Kuckreja , Parul Gupta , Injy Hamed , Thamar Solorio , Muhammad Haris Khan , Abhinav Dhall

Social media platforms and online streaming services have spawned a new breed of Hate Speech (HS). Due to the massive amount of user-generated content on these sites, modern machine learning techniques are found to be feasible and…

Computation and Language · Computer Science 2022-06-02 Nauros Romim , Mosahed Ahmed , Md. Saiful Islam , Arnab Sen Sharma , Hriteshwar Talukder , Mohammad Ruhul Amin

The rapid advancement of deep generative models has significantly improved the realism of synthetic media, presenting both opportunities and security challenges. While deepfake technology has valuable applications in entertainment and…

Machine Learning · Computer Science 2025-06-09 Arnesh Batra , Anushk Kumar , Jashn Khemani , Arush Gumber , Arhan Jain , Somil Gupta

Recent advances in Text-to-Speech (TTS) systems have substantially increased the realism of synthetic speech, raising new challenges for audio deepfake detection. This work presents a comparative evaluation of three state-of-the-art TTS…

With the huge technological advances introduced by deep learning in audio & speech processing, many novel synthetic speech techniques achieved incredible realistic results. As these methods generate realistic fake human voices, they can be…

Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets…