中文
相关论文

相关论文: An Efficient Temporary Deepfake Location Approach …

200 篇论文

Spoof diarization identifies ``what spoofed when" in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward this task, prior work proposed a two-branch model for…

音频与语音处理 · 电气工程与系统科学 2025-09-17 Kyo-Won Koo , Chan-yeong Lim , Jee-weon Jung , Hye-jin Shim , Ha-Jin Yu

With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake…

声音 · 计算机科学 2025-06-04 Chi Ding , Junxiao Xue , Cong Wang , Hao Zhou

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker…

声音 · 计算机科学 2024-12-25 Xuechen Liu , Junichi Yamagishi , Md Sahidullah , Tomi kinnunen

Better generative models and larger datasets have led to more realistic fake videos that can fool the human eye but produce temporal and spatial artifacts that deep learning approaches can detect. Most current Deepfake detection methods…

计算机视觉与模式识别 · 计算机科学 2020-06-29 Oscar de Lima , Sean Franklin , Shreshtha Basu , Blake Karwoski , Annet George

When the task of locating manipulation regions in partially-fake audio (PFA) involves cross-domain datasets, the performance of deep learning models drops significantly due to the shift between the source and target domains. To address this…

声音 · 计算机科学 2024-07-12 Siding Zeng , Jiangyan Yi , Jianhua Tao , Yujie Chen , Shan Liang , Yong Ren , Xiaohui Zhang

Synthetic voice and splicing audio clips have been generated to spoof Internet users and artificial intelligence (AI) technologies such as voice authentication. Existing research work treats spoofing countermeasures as a binary…

音频与语音处理 · 电气工程与系统科学 2022-11-30 Lei Wang , Benedict Yeoh , Jun Wah Ng

Recent advances in voice cloning and lip synchronization models have enabled Synthesized Audiovisual Forgeries (SAVFs), where both audio and visuals are manipulated to mimic a target speaker. This significantly increases the risk of…

声音 · 计算机科学 2025-07-18 Minyoung Kim , Sehwan Park , Sungmin Cha , Paul Hongsuck Seo

The spread of Deepfake videos has caused a trust crisis and impaired social stability. Although numerous approaches have been proposed to address the challenges of Deepfake detection and localization, there is still a lack of systematic…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Wenbo Xu , Wei Lu , Xiangyang Luo

Recent multimodal deepfake detection methods designed for generalization conjecture that single-stage supervised training struggles to generalize across unseen manipulations and datasets. However, such approaches that target generalization…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Ashutosh Anshul , Shreyas Gopal , Deepu Rajan , Eng Siong Chng

Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest…

Speaker verification systems have been used in many production scenarios in recent years. Unfortunately, they are still highly prone to different kinds of spoofing attacks such as voice conversion and speech synthesis, etc. In this paper,…

音频与语音处理 · 电气工程与系统科学 2021-09-07 Junxiao Xue , Hao Zhou , Yabo Wang

Spoofing-robust automatic speaker verification (SASV) systems are a crucial technology for the protection against spoofed speech. In this study, we focus on logical access attacks and introduce a novel approach to SASV tasks. A novel…

音频与语音处理 · 电气工程与系统科学 2024-12-24 Avishai Weizman , Yehuda Ben-Shimol , Itshak Lapidot

Copy-move forgery on speech (CMF), coupled with post-processing techniques, presents a great challenge to the forensic detection and localization of tampered areas. Most of the existing CMF detection approaches necessitate pre-segmentation…

音频与语音处理 · 电气工程与系统科学 2023-09-11 Dong Yang , Mingle Liu , Muyong Cao

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

声音 · 计算机科学 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio…

音频与语音处理 · 电气工程与系统科学 2020-11-12 Tyler Vuong , Yangyang Xia , Richard Stern

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a…

声音 · 计算机科学 2024-09-11 Tuan Duy Nguyen Le , Kah Kuan Teh , Huy Dat Tran

We propose a transfer deep learning (TDL) framework that can transfer the knowledge obtained from a single-modal neural network to a network with a different modality. Specifically, we show that we can leverage speech data to fine-tune the…

神经与进化计算 · 计算机科学 2016-02-19 Seungwhan Moon , Suyoun Kim , Haohan Wang

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines…

音频与语音处理 · 电气工程与系统科学 2025-07-08 Davide Berghi , Philip J. B. Jackson

The state-of-art models for speech synthesis and voice conversion are capable of generating synthetic speech that is perceptually indistinguishable from bonafide human speech. These methods represent a threat to the automatic speaker…

机器学习 · 计算机科学 2019-07-11 Moustafa Alzantot , Ziqi Wang , Mani B. Srivastava