中文
相关论文

相关论文: Proactive Detection of Voice Cloning with Localize…

200 篇论文

In the realm of audio watermarking, it is challenging to simultaneously encode imperceptible messages while enhancing the message capacity and robustness. Although recent advancements in deep learning-based methods bolster the message…

声音 · 计算机科学 2024-11-05 Mayank Kumar Singh , Naoya Takahashi , Weihsiang Liao , Yuki Mitsufuji

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leveraging the synergy…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Yachao Liang , Min Yu , Gang Li , Jianguo Jiang , Boquan Li , Feng Yu , Ning Zhang , Xiang Meng , Weiqing Huang

Neural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often…

This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic…

声音 · 计算机科学 2025-07-16 Andrew Valdivia , Yueming Zhang , Hailu Xu , Amir Ghasemkhani , Xin Qin

Watermarking has emerged as a promising technique to track AI-generated content and differentiate it from authentic human creations. While prior work extensively studies watermarking for autoregressive large language models (LLMs) and image…

密码学与安全 · 计算机科学 2026-02-16 Avi Bagchi , Akhil Bhimaraju , Moulik Choraria , Daniel Alabi , Lav R. Varshney

Deepfake speech attribution remains challenging for existing solutions. Classifier-based solutions often fail to generalize to domain-shifted samples, and watermarking-based solutions are easily compromised by distortions like codec…

音频与语音处理 · 电气工程与系统科学 2025-10-16 Wanying Ge , Xin Wang , Junichi Yamagishi

In this paper, we propose a technique to alleviate the quality degradation caused by collapsed speech segments sometimes generated by the WaveNet vocoder. The effectiveness of the WaveNet vocoder for generating natural speech from acoustic…

音频与语音处理 · 电气工程与系统科学 2018-08-10 Yi-Chiao Wu , Kazuhiro Kobayashi , Tomoki Hayashi , Patrick Lumban Tobing , Tomoki Toda

Proprietary large language models (LLMs) face risks of intellectual property (IP) violation, as adversaries can replicate an LLM by collecting input-output pairs to train a surrogate model, causing financial setbacks. Watermarks offer a…

密码学与安全 · 计算机科学 2026-05-25 Kieu Dang , Phung Lai , NhatHai Phan , Yelong Shen , Ruoming Jin

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

声音 · 计算机科学 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan

Advances in AI technology have made voice cloning increasingly accessible, leading to a rise in fraud involving AI-generated audio forgeries. This highlights the need to covertly embed information and verify the authenticity and integrity…

密码学与安全 · 计算机科学 2024-08-28 Guang Yang

With the advancement of AIGC technologies, the modalities generated by models have expanded from images and videos to 3D objects, leading to an increasing number of works focused on 3D Gaussian Splatting (3DGS) generative models. Existing…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Runyi Li , Xuanyu Zhang , Chuhan Tong , Zhipei Xu , Jian Zhang

As artificial intelligence surpasses human capabilities in text generation, the necessity to authenticate the origins of AI-generated content has become paramount. Unbiased watermarks offer a powerful solution by embedding statistical…

计算与语言 · 计算机科学 2025-08-07 Ruibo Chen , Yihan Wu , Junfeng Guo , Heng Huang

As an effective method for intellectual property (IP) protection, model watermarking technology has been applied on a wide variety of deep neural networks (DNN), including speech classification models. However, how to design a black-box…

声音 · 计算机科学 2022-05-03 Haozhe Chen , Weiming Zhang , Kunlin Liu , Kejiang Chen , Han Fang , Nenghai Yu

We propose a novel audio watermarking system that is robust to the distortion due to the indoor acoustic propagation channel between the loudspeaker and the receiving microphone. The system utilizes a set of new algorithms that effectively…

多媒体 · 计算机科学 2019-03-21 Yuan-Yen Tai , Mohamed F. Mansour

Most image watermarking systems focus on robustness, capacity, and imperceptibility while treating the embedded payload as meaningless bits. This bit-centric view imposes a hard ceiling on capacity and prevents watermarks from carrying…

密码学与安全 · 计算机科学 2025-10-02 Gautier Evennou , Vivien Chappelier , Ewa Kijak

Audio grounding, or speech-driven open-set object detection, aims to localize and identify objects directly from speech, enabling generalization beyond predefined categories. This task is crucial for applications like human-robot…

声音 · 计算机科学 2025-09-23 Wenhuan Lu , Xinyue Song , Wenjun Ke , Zhizhi Yu , Wenhao Yang , Jianguo Wei

The advancement of artificial intelligence generated content (AIGC) has created a pressing need for robust image watermarking that can withstand both conventional signal processing and novel semantic editing attacks. Current deep…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Yichao Tang , Mingyang Li , Di Miao , Sheng Li , Zhenxing Qian , Xinpeng Zhang

Geo-localization aims to infer the geographic origin of a given signal. In computer vision, geo-localization has served as a demanding benchmark for compositional reasoning and is relevant to public safety. In contrast, progress on audio…

声音 · 计算机科学 2026-01-07 Ruixing Zhang , Zihan Liu , Leilei Sun , Tongyu Zhu , Weifeng Lv

The performance of sound event detection methods can significantly degrade when they are used in unseen conditions (e.g. recording devices, ambient noise). Domain adaptation is a promising way to tackle this problem. In this paper, we…

声音 · 计算机科学 2019-11-26 Shayan Gharib , Konstantinos Drossos , Eemi Fagerlund , Tuomas Virtanen

Recent progress in generative AI technology has made audio deepfakes remarkably more realistic. While current research on anti-spoofing systems primarily focuses on assessing whether a given audio sample is fake or genuine, there has been…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Nicholas Klein , Tianxiang Chen , Hemlata Tak , Ricardo Casal , Elie Khoury