中文
相关论文

相关论文: Generalized Spoofing Detection Inspired from Audio…

200 篇论文

Spoofing detection is today a mainstream research topic. Standard metrics can be applied to evaluate the performance of isolated spoofing detection solutions and others have been proposed to support their evaluation when they are combined…

音频与语音处理 · 电气工程与系统科学 2025-04-17 Hye-jin Shim , Jee-weon Jung , Tomi Kinnunen , Nicholas Evans , Jean-Francois Bonastre , Itshak Lapidot

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps. While long-range dependencies are difficult to model directly in the time domain, we show that they can…

音频与语音处理 · 电气工程与系统科学 2019-06-05 Sean Vasquez , Mike Lewis

Deepfakes represent one of the toughest challenges in the world of Cybersecurity and Digital Forensics, especially considering the high-quality results obtained with recent generative AI-based solutions. Almost all generative models leave…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Orazio Pontorno , Luca Guarnera , Sebastiano Battiato

Due to the successful development of deep image generation technology, visual data forgery detection would play a more important role in social and economic security. Existing forgery detection methods suffer from unsatisfactory…

计算机视觉与模式识别 · 计算机科学 2023-07-24 Decheng Liu , Tao Chen , Chunlei Peng , Nannan Wang , Ruimin Hu , Xinbo Gao

Recent advances in neural audio codec-based speech generation (CoSG) models have produced remarkably realistic audio deepfakes. We refer to deepfake speech generated by CoSG systems as codec-based deepfake, or CodecFake. Although existing…

声音 · 计算机科学 2025-08-05 Xuanjun Chen , I-Ming Lin , Lin Zhang , Jiawei Du , Haibin Wu , Hung-yi Lee , Jyh-Shing Roger Jang

Synthetic facial videos have proliferated across social media faster than platform moderation can respond, raising the cost of disinformation and identity-based attacks. Frame-level deepfake detectors degrade sharply as generator quality…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Mohammadreza Rashidi , Raja Hashim Ali , Sami Ur Rahman

We show how replay attacks undermine audio deepfake detection: By playing and re-recording deepfake audio through various speakers and microphones, we make spoofed samples appear authentic to the detection model. To study this phenomenon in…

In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context…

音频与语音处理 · 电气工程与系统科学 2020-09-10 Wei Xia , John H. L. Hansen

Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio…

音频与语音处理 · 电气工程与系统科学 2020-11-12 Tyler Vuong , Yangyang Xia , Richard Stern

Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed…

音频与语音处理 · 电气工程与系统科学 2025-05-21 Taewoo Kim , Guisik Kim , Choongsang Cho , Young Han Lee

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. While these…

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental sounds have…

声音 · 计算机科学 2025-09-30 Han Yin , Yang Xiao , Rohan Kumar Das , Jisheng Bai , Haohe Liu , Wenwu Wang , Mark D Plumbley

Spoof detectors are classifiers that are trained to distinguish spoof fingerprints from bonafide ones. However, state of the art spoof detectors do not generalize well on unseen spoof materials. This study proposes a style transfer based…

计算机视觉与模式识别 · 计算机科学 2019-12-10 Rohit Gajawada , Additya Popli , Tarang Chugh , Anoop Namboodiri , Anil K. Jain

This paper presents a system for detecting fake audio-visual content (i.e., video deepfake), developed for Track 2 of the DDL Challenge. The proposed system employs a two-stage framework, comprising unimodal detection and multimodal score…

多媒体 · 计算机科学 2026-02-03 Qingcao Li , Miao He , Liang Yi , Qing Wen , Yitao Zhang , Hongshuo Jin , Peng Cheng , Zhongjie Ba , Li Lu , Kui Ren

Environmental audio tagging is a newly proposed task to predict the presence or absence of a specific audio event in a chunk. Deep neural network (DNN) based methods have been successfully adopted for predicting the audio tags in the…

声音 · 计算机科学 2017-02-28 Yong Xu , Qiuqiang Kong , Qiang Huang , Wenwu Wang , Mark D. Plumbley

This paper describes the USTC-KXDIGIT system submitted to the ASVspoof5 Challenge for Track 1 (speech deepfake detection) and Track 2 (spoofing-robust automatic speaker verification, SASV). Track 1 showcases a diverse range of technical…

Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically…

声音 · 计算机科学 2026-02-19 Yixuan Xiao , Florian Lux , Alejandro Pérez-González-de-Martos , Ngoc Thang Vu

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two rounds of evaluation in Track 3.2 with a similar training…

音频与语音处理 · 电气工程与系统科学 2022-04-21 Rui Yan , Cheng Wen , Shuran Zhou , Tingwei Guo , Wei Zou , Xiangang Li

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms…