English
Related papers

Related papers: VoiceCloak: A Multi-Dimensional Defense Framework …

200 papers

We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models relying on language backbones have achieved remarkable…

Computation and Language · Computer Science 2025-06-03 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Ruibo Fu , Wei Liang , Dong Yu

Diffusion models (DMs) are regarded as one of the most advanced generative models today, yet recent studies suggest that they are vulnerable to backdoor attacks, which establish hidden associations between particular input patterns and…

Cryptography and Security · Computer Science 2024-08-23 Jiang Hao , Xiao Jin , Hu Xiaoguang , Chen Tianyou , Zhao Jiajia

Identity, accent, style, and emotions are essential components of human speech. Voice conversion (VC) techniques process the speech signals of two input speakers and other modalities of auxiliary information such as prompts and emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-09 Xining Song , Zhihua Wei , Rui Wang , Haixiao Hu , Yanxiang Chen , Meng Han

Recent advances in diffusion models have introduced a new era of text-guided image manipulation, enabling users to create realistic edited images with simple textual prompts. However, there is significant concern about the potential misuse…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 June Suk Choi , Kyungmin Lee , Jongheon Jeong , Saining Xie , Jinwoo Shin , Kimin Lee

Pretrained language models have significantly advanced performance across various natural language processing tasks. However, adversarial attacks continue to pose a critical challenge to systems built using these models, as they can be…

Computation and Language · Computer Science 2025-05-20 Zhenhao Li , Huichi Zhou , Marek Rei , Lucia Specia

Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC). Conventional speech representation learning methods in VC only factorize speech as speaker and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-06 Jie Wang , Jingbei Li , Xintao Zhao , Zhiyong Wu , Shiyin Kang , Helen Meng

Vision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, posing serious risks. Existing medical attacks focus on secondary objectives such as model…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Akash Ghosh , Subhadip Baidya , Sriparna Saha , Xiuying Chen

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we…

Sound · Computer Science 2026-01-05 Akanksha Chuchra , Shukesh Reddy , Sudeepta Mishra , Abhijit Das , Abhinav Dhall

Generalization is a main issue for current audio deepfake detectors, which struggle to provide reliable results on out-of-distribution data. Given the speed at which more and more accurate synthesis methods are developed, it is very…

Sound · Computer Science 2024-07-02 Alessandro Pianese , Davide Cozzolino , Giovanni Poggi , Luisa Verdoliva

Voice conversion (VC) consists of digitally altering the voice of an individual to manipulate part of its content, primarily its identity, while maintaining the rest unchanged. Research in neural VC has accomplished considerable…

Sound · Computer Science 2021-07-28 Laurent Benaroya , Nicolas Obin , Axel Roebel

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive…

Sound · Computer Science 2025-06-02 Kaidi Wang , Wenhao Guan , Ziyue Jiang , Hukai Huang , Peijie Chen , Weijie Wu , Qingyang Hong , Lin Li

Voice authentication has undergone significant changes from traditional systems that relied on handcrafted acoustic features to deep learning models that can extract robust speaker embeddings. This advancement has expanded its applications…

Cryptography and Security · Computer Science 2025-09-17 Kamel Kamel , Keshav Sood , Hridoy Sankar Dutta , Sunil Aryal

With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices. Existing defenses…

Machine Learning · Computer Science 2025-05-20 Seungmin Kim , Sohee Park , Donghyun Kim , Jisu Lee , Daeseon Choi

Although voice conversion (VC) systems have shown a remarkable ability to transfer voice style, existing methods still have an inaccurate pitch and low speaker adaptation quality. To address these challenges, we introduce Diff-HierVC, a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

Despite the success of diffusion-based customization methods on visual content creation, increasing concerns have been raised about such techniques from both privacy and political perspectives. To tackle this issue, several…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Feifei Wang , Zhentao Tan , Tianyi Wei , Yue Wu , Qidong Huang

With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake…

Sound · Computer Science 2025-06-04 Chi Ding , Junxiao Xue , Cong Wang , Hao Zhou

Voice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective through streaming model…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-19 Ziqian Ning , Yuepeng Jiang , Pengcheng Zhu , Shuai Wang , Jixun Yao , Lei Xie , Mengxiao Bi

Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Michael Neri , Tuomas Virtanen

Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinformation. Existing protection methods are largely limited to a single modality, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Wenli Zhang , Xianglong Shi , Sirui Zhao , Xinqi Chen , Guo Cheng , Yifan Xu , Tong Xu , Yong Liao

Voice authentication has become an integral part in security-critical operations, such as bank transactions and call center conversations. The vulnerability of automatic speaker verification systems (ASVs) to spoofing attacks instigated the…

Cryptography and Security · Computer Science 2021-08-02 Andre Kassis , Urs Hengartner