中文
相关论文

相关论文: Proactive Detection of Voice Cloning with Localize…

200 篇论文

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of…

声音 · 计算机科学 2023-10-18 Yu Chen , Xinyuan Qian , Zexu Pan , Kainan Chen , Haizhou Li

With the rapid advancement of speech generative models, unauthorized voice cloning poses significant privacy and security risks. Speech watermarking offers a viable solution for tracing sources and preventing misuse. Current watermarking…

声音 · 计算机科学 2025-09-30 Yang Cui , Peter Pan , Lei He , Sheng Zhao

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread…

多媒体 · 计算机科学 2025-06-09 Marcel Klemt , Carlotta Segna , Anna Rohrbach

Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like…

Spoofed audio, i.e. audio that is manipulated or AI-generated deepfake audio, is difficult to detect when only using acoustic features. Some recent innovative work involving AI-spoofed audio detection models augmented with phonetic and…

声音 · 计算机科学 2024-10-22 Zahra Khanjani , Christine Mallinson , James Foulds , Vandana P Janeja

Automated Audio Captioning is a multimodal task that aims to convert audio content into natural language. The assessment of audio captioning systems is typically based on quantitative metrics applied to text data. Previous studies have…

声音 · 计算机科学 2024-03-28 Gijs Wijngaard , Elia Formisano , Bruno L. Giordano , Michel Dumontier

Text recognition in the wild is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest vision and language processing are effective for scene text recognition. Yet, solving edit errors such as…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Wenwen Yu , Mingyu Liu , Biao Yang , Enming Zhang , Deqiang Jiang , Xing Sun , Yuliang Liu , Xiang Bai

Audio watermarking has been widely applied in copyright protection and source tracing. However, due to the inherent characteristics of audio signals, watermark localization and resistance to desynchronization attacks remain significant…

密码学与安全 · 计算机科学 2025-09-03 Zhenliang Gan , Xiaoxiao Hu , Sheng Li , Zhenxing Qian , Xinpeng Zhang

Invisible watermarking is essential for tracing the provenance of digital content. However, training state-of-the-art models remains notoriously difficult, with current approaches often struggling to balance robustness against true…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Tomáš Souček , Pierre Fernandez , Hady Elsahar , Sylvestre-Alvise Rebuffi , Valeriu Lacatusu , Tuan Tran , Tom Sander , Alexandre Mourachko

Large language models generate high-quality responses with potential misinformation, underscoring the need for regulation by distinguishing AI-generated and human-written texts. Watermarking is pivotal in this context, which involves…

机器学习 · 计算机科学 2024-06-07 Mingjia Huo , Sai Ashish Somayajula , Youwei Liang , Ruisi Zhang , Farinaz Koushanfar , Pengtao Xie

Passive acoustic monitoring offers a scalable, non-invasive method for tracking global biodiversity and anthropogenic impacts on species. Although deep learning has become a vital tool for processing this data, current models are…

机器学习 · 计算机科学 2023-08-10 David Robinson , Adelaide Robinson , Lily Akrapongpisak

Digital audio watermarking consists in inserting a message into audio signals in a transparent way and can be used to allow automatic recognition of audio material and management of the copyrights. We propose a perceptual loss function to…

音频与语音处理 · 电气工程与系统科学 2024-11-05 Martin Moritz , Toni Olán , Tuomas Virtanen

Nowadays voice search for points of interest (POI) is becoming increasingly popular. However, speech recognition for local POI has remained to be a challenge due to multi-dialect and massive POI. This paper improves speech recognition…

音频与语音处理 · 电气工程与系统科学 2021-07-08 Songjun Cao , Yike Zhang , Xiaobing Feng , Long Ma

Earables (ear wearables) is rapidly emerging as a new platform encompassing a diverse range of personal applications. The traditional authentication methods hence become less applicable and inconvenient for earables due to their limited…

密码学与安全 · 计算机科学 2022-04-18 Zi Wang , Jie Yang

Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Hoang-Son Vo , Quang-Vinh Nguyen , Seungwon Kim , Hyung-Jeong Yang , Soonja Yeom , Soo-Hyung Kim

Acoustic scene classification (ASC) predominantly relies on supervised approaches. However, acquiring labeled data for training ASC models is often costly and time-consuming. Recently, self-supervised learning (SSL) has emerged as a…

声音 · 计算机科学 2024-08-28 Yiqiang Cai , Shengchen Li , Xi Shao

The recent rise in capabilities of AI-based music generation tools has created an upheaval in the music industry, necessitating the creation of accurate methods to detect such AI-generated content. This can be done using audio-based…

The growing threat of deepfakes and manipulated media necessitates a radical rethinking of media authentication. Existing methods for watermarking synthetic data fall short, as they can be easily removed or altered, and current deepfake…

密码学与安全 · 计算机科学 2024-11-27 Bhaktipriya Radharapu , Harish Krishna

The rapid proliferation of AI-manipulated or generated audio deepfakes poses serious challenges to media integrity and election security. Current AI-driven detection solutions lack explainability and underperform in real-world settings. In…

机器学习 · 计算机科学 2024-10-11 Georgia Channing , Juil Sock , Ronald Clark , Philip Torr , Christian Schroeder de Witt

The rapid advancement of generative models has led to the synthesis of real-fake ambiguous voices. To erase the ambiguity, embedding watermarks into the frequency-domain features of synthesized voices has become a common routine. However,…

密码学与安全 · 计算机科学 2025-06-24 Yue Li , Weizhi Liu , Dongdong Lin , Hui Tian , Hongxia Wang