中文
相关论文

相关论文: Harder or Different? Understanding Generalization …

200 篇论文

Automatic speaker verification (ASV) systems utilize the biometric information in human speech to verify the speaker's identity. The techniques used for performing speaker verification are often vulnerable to malicious attacks that attempt…

声音 · 计算机科学 2020-11-26 Yang Gao , Jiachen Lian , Bhiksha Raj , Rita Singh

State-of-the-art methods for audio generation suffer from fingerprint artifacts and repeated inconsistencies across temporal and spectral domains. Such artifacts could be well captured by the frequency domain analysis over the spectrogram.…

声音 · 计算机科学 2021-06-29 Yang Gao , Tyler Vuong , Mahsa Elyasi , Gaurav Bharaj , Rita Singh

In this work, we consider the task of automated emphasis detection for spoken language. This problem is challenging in that emphasis is affected by the particularities of speech of the subject, for example the subject accent, dialect or…

机器学习 · 计算机科学 2023-05-16 Eran Kaufman , Lee-Ad Gottlieb

AI-synthesized speech, also known as deepfake speech, has recently raised significant concerns due to the rapid advancement of speech synthesis and speech conversion techniques. Previous works often rely on distinguishing synthesizer…

声音 · 计算机科学 2024-11-15 Kuiyuan Zhang , Zhongyun Hua , Yushu Zhang , Yifang Guo , Tao Xiang

Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers…

声音 · 计算机科学 2025-09-12 Chin Yuen Kwok , Jia Qi Yip , Zhen Qiu , Chi Hung Chi , Kwok Yan Lam

Deepfakes, created using advanced AI techniques such as Variational Autoencoder and Generative Adversarial Networks, have evolved from research and entertainment applications into tools for malicious activities, posing significant threats…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Yamini Sri Krubha , Aryana Hou , Braden Vester , Web Walker , Xin Wang , Li Lin , Shu Hu

Existing deepfake speech detection systems lack generalizability to unseen attacks (i.e., samples generated by generative algorithms not seen during training). Recent studies have explored the use of universal speech representations to…

声音 · 计算机科学 2023-09-18 Yi Zhu , Saurabh Powar , Tiago H. Falk

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

声音 · 计算机科学 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

Deepfakes powered by advanced machine learning models present a significant and evolving threat to identity verification and the authenticity of digital media. Although numerous detectors have been developed to address this problem, their…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Viacheslav Pirogov , Maksim Artemev

This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light…

音频与语音处理 · 电气工程与系统科学 2022-10-14 Yuxiang Zhang , Jingze Lu , Xingming Wang , Zhuo Li , Runqiu Xiao , Wenchao Wang , Ming Li , Pengyuan Zhang

Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious attackers may exploit these technologies, causing misuse…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Haibin Wu , Yuan Tseng , Hung-yi Lee

Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deepfake detectors. However, recent studies have shown that the current audio deepfake models fall short of this desideratum. In…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Octavian Pascu , Adriana Stan , Dan Oneata , Elisabeta Oneata , Horia Cucu

Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment…

声音 · 计算机科学 2023-01-27 Jianwei Zhang , Julie Liss , Suren Jayasuriya , Visar Berisha

The rapid proliferation of AI-generated content, driven by advances in generative adversarial networks, diffusion models, and multimodal large language models, has made the creation and dissemination of synthetic media effortless,…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Guangyu Lin , Li Lin , Christina P. Walker , Daniel S. Schiff , Shu Hu

Deepfakes are synthetically generated media often devised with malicious intent. They have become increasingly more convincing with large training datasets advanced neural networks. These fakes are readily being misused for slander,…

密码学与安全 · 计算机科学 2022-03-30 Nicolas M. Müller , Franziska Dieckmann , Jennifer Williams

Deepfake detection faces a critical generalization hurdle, with performance deteriorating when there is a mismatch between the distributions of training and testing data. A broadly received explanation is the tendency of these detectors to…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Zhiyuan Yan , Yuhao Luo , Siwei Lyu , Qingshan Liu , Baoyuan Wu

Recent advances in deep learning have led to substantial improvements in deepfake generation, resulting in fake media with a more realistic appearance. Although deepfake media have potential application in a wide range of areas and are…

计算机视觉与模式识别 · 计算机科学 2022-02-15 Trung-Nghia Le , Huy H Nguyen , Junichi Yamagishi , Isao Echizen

Although existing face anti-spoofing (FAS) methods achieve high accuracy in intra-domain experiments, their effects drop severely in cross-domain scenarios because of poor generalization. Recently, multifarious techniques have been…

计算机视觉与模式识别 · 计算机科学 2022-01-03 Shice Liu , Shitao Lu , Hongyi Xu , Jing Yang , Shouhong Ding , Lizhuang Ma

Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using…

声音 · 计算机科学 2025-08-19 Xuanjun Chen , I-Ming Lin , Lin Zhang , Haibin Wu , Hung-yi Lee , Jyh-Shing Roger Jang

Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically…

声音 · 计算机科学 2026-02-19 Yixuan Xiao , Florian Lux , Alejandro Pérez-González-de-Martos , Ngoc Thang Vu