中文
相关论文

相关论文: RTCFake: Speech Deepfake Detection in Real-Time Co…

200 篇论文

Multi-frame algorithms for single-channel speech enhancement are able to take advantage from short-time correlations within the speech signal. Deep Filtering (DF) was proposed to directly estimate a complex filter in frequency domain to…

音频与语音处理 · 电气工程与系统科学 2023-05-16 Hendrik Schröter , Tobias Rosenkranz , Alberto N. Escalante-B. , Andreas Maier

Audio deepfake detection has become increasingly challenging due to rapid advances in speech synthesis and voice conversion technologies, particularly under channel distortions, replay attacks, and real-world recording conditions. This…

音频与语音处理 · 电气工程与系统科学 2026-01-13 K. A. Shahriar

This research explores the positive application of deepfake technology for upper body generation, specifically sign language for the Deaf and Hard of Hearing (DHoH) community. Given the complexity of sign language and the scarcity of…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Shahzeb Naeem , Muhammad Riyyan Khan , Usman Tariq , Abhinav Dhall , Carlos Ivan Colon , Hasan Al-Nashash

Disentangled representation learning in speech processing has lagged behind other domains, largely due to the lack of datasets with annotated generative factors for robust evaluation. To address this, we propose SynSpeech, a novel…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Yusuf Brima , Ulf Krumnack , Simone Pika , Gunther Heidemann

Recent advances in generative models for language have enabled the creation of convincing synthetic text or deepfake text. Prior work has demonstrated the potential for misuse of deepfake text to mislead content consumers. Therefore,…

In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier…

声音 · 计算机科学 2024-07-03 Lam Pham , Phat Lam , Truong Nguyen , Huyen Nguyen , Alexander Schindler

The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive detection methods, we propose StreamMark, a novel deep…

音频与语音处理 · 电气工程与系统科学 2026-04-15 Zhentao Liu , Milos Cernak

Self-supervised speech models learn representations that capture both content and speaker information. Yet this entanglement creates problems: content tasks suffer from speaker bias, and privacy concerns arise when speaker identity leaks…

声音 · 计算机科学 2026-04-02 Xiaoxu Zhu , Junhua Li , Aaron J. Li , Guangchao Yao , Xiaojie Yu

Humans use context to assess the veracity of information. However, current audio deepfake detectors only analyze the audio file without considering either context or transcripts. We create and analyze a Journalist-provided Deepfake Dataset…

Deep learning has enabled realistic face manipulation (i.e., deepfake), which poses significant concerns over the integrity of the media in circulation. Most existing deep learning techniques for deepfake detection can achieve promising…

计算机视觉与模式识别 · 计算机科学 2022-12-26 Bosheng Yan , Chang-Tsun Li , Xuequan Lu

Good datasets are essential for developing and benchmarking any machine learning system. Their importance is even more extreme for safety critical applications such as deepfake detection - the focus of this paper. Here we reveal that two of…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Stefan Smeu , Dragos-Alexandru Boldisor , Dan Oneata , Elisabeta Oneata

Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current…

Modern deepfakes have evolved into localized and intermittent manipulations that require fine-grained temporal localization to mitigate severe digital security risks. The prohibitive cost of frame-level annotation makes weakly supervised…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Midou Guo , Qilin Yin , Wei Lu , Rui Yang

In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic and pose a greater…

声音 · 计算机科学 2024-08-08 Vinaya Sree Katamneni , Ajita Rattani

Real-time deepfake, a type of generative AI, is capable of "creating" non-existing contents (e.g., swapping one's face with another) in a video. It has been, very unfortunately, misused to produce deepfake videos (during web conferences,…

计算机视觉与模式识别 · 计算机科学 2024-09-18 Zhixin Xie , Jun Luo

Deepfakes offer great potential for innovation and creativity, but they also pose significant risks to privacy, trust, and security. With a vast Hindi-speaking population, India is particularly vulnerable to deepfake-driven misinformation…

声音 · 计算机科学 2024-11-26 Sukhandeep Kaur , Mubashir Buhari , Naman Khandelwal , Priyansh Tyagi , Kiran Sharma

Fake content has grown at an incredible rate over the past few years. The spread of social media and online platforms makes their dissemination on a large scale increasingly accessible by malicious actors. In parallel, due to the growing…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Luca Maiano , Lorenzo Papa , Ketbjano Vocaj , Irene Amerini

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental sounds have…

声音 · 计算机科学 2025-09-30 Han Yin , Yang Xiao , Rohan Kumar Das , Jisheng Bai , Haohe Liu , Wenwu Wang , Mark D Plumbley

The emergence of text-to-image generative models has revolutionized the field of deepfakes, enabling the creation of realistic and convincing visual content directly from textual descriptions. However, this advancement presents considerably…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Yabin Wang , Zhiwu Huang , Zhiheng Ma , Xiaopeng Hong

Recent advances in technology for hyper-realistic visual and audio effects provoke the concern that deepfake videos of political speeches will soon be indistinguishable from authentic video recordings. The conventional wisdom in…

人机交互 · 计算机科学 2024-01-17 Matthew Groh , Aruna Sankaranarayanan , Nikhil Singh , Dong Young Kim , Andrew Lippman , Rosalind Picard