English
Related papers

Related papers: Adversarial speech for voice privacy protection fr…

200 papers

With rapid globalization, the need to build inclusive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

We present a speaker conditioned text-to-speech (TTS) system aimed at addressing challenges in generating speech for unseen speakers and supporting diverse Indian languages. Our method leverages a diffusion-based TTS architecture, where a…

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Cheng Wen , Tingwei Guo , Xingjun Tan , Rui Yan , Shuran Zhou , Chuandong Xie , Wei Zou , Xiangang Li

High-performance spoofing countermeasure systems for automatic speaker verification (ASV) have been proposed in the ASVspoof 2019 challenge. However, the robustness of such systems under adversarial attacks has not been studied yet. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-22 Songxiang Liu , Haibin Wu , Hung-yi Lee , Helen Meng

Text embedding inversion attacks reconstruct original sentences from latent representations, posing severe privacy threats in collaborative inference and edge computing. We propose TextCrafter, an optimization-based adversarial perturbation…

Cryptography and Security · Computer Science 2026-01-23 Duoxun Tang , Xinhang Jiang , Jiajun Niu

Lack of moderation in online communities enables participants to incur in personal aggression, harassment or cyberbullying, issues that have been accentuated by extremist radicalisation in the contemporary post-truth politics scenario. This…

Computation and Language · Computer Science 2018-01-08 Nestor Rodriguez , Sergio Rojas-Galeano

Recent breakthroughs in text-to-speech (TTS) voice cloning have raised serious privacy concerns, allowing highly accurate vocal identity replication from just a few seconds of reference audio, while retaining the speaker's vocal…

Sound · Computer Science 2025-05-27 Renyuan Li , Zhibo Liang , Haichuan Zhang , Tianyu Shi , Zhiyuan Cheng , Jia Shi , Carl Yang , Mingjie Tang

Large audio-language models increasingly operate on raw speech inputs, enabling more seamless integration across domains such as voice assistants, education, and clinical triage. This transition, however, introduces a distinct class of…

Computation and Language · Computer Science 2026-02-02 Ye Yu , Haibo Jin , Yaoning Yu , Jun Zhuang , Haohan Wang

In the past few years, it has been shown that deep learning systems are highly vulnerable under attacks with adversarial examples. Neural-network-based automatic speech recognition (ASR) systems are no exception. Targeted and untargeted…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-07 Matías Pizarro , Dorothea Kolossa , Asja Fischer

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity…

Sound · Computer Science 2022-08-09 Huaizhen Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Zhen Zeng , Edward Xiao , Jing Xiao

As language models become increasingly integrated into our digital lives, Personalized Text Generation (PTG) has emerged as a pivotal component with a wide range of applications. However, the bias inherent in user written text, often used…

Computation and Language · Computer Science 2023-10-24 Nan Wang , Qifan Wang , Yi-Chia Wang , Maziar Sanjabi , Jingzhou Liu , Hamed Firooz , Hongning Wang , Shaoliang Nie

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech synthesis technology…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-23 Dongyang Dai , Li Chen , Yuping Wang , Mu Wang , Rui Xia , Xuchen Song , Zhiyong Wu , Yuxuan Wang

Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and digital signal processing-based watermarking methods have…

Sound · Computer Science 2026-03-17 Yigitcan Özer , Wanying Ge , Zhe Zhang , Xin Wang , Junichi Yamagishi

Voice deepfake attacks, which artificially impersonate human speech for malicious purposes, have emerged as a severe threat. Existing defenses typically inject noise into human speech to compromise voice encoders in speech synthesis models.…

Sound · Computer Science 2025-08-26 Yuanda Wang , Bocheng Chen , Hanqing Guo , Guangjing Wang , Weikang Ding , Qiben Yan

We construct targeted audio adversarial examples on automatic speech recognition. Given any audio waveform, we can produce another that is over 99.9% similar, but transcribes as any phrase we choose (recognizing up to 50 characters per…

Machine Learning · Computer Science 2018-04-02 Nicholas Carlini , David Wagner

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

Given the speech generation framework that represents the speaker attribute with an embedding vector, asynchronous voice anonymization can be achieved by modifying the speaker embedding derived from the original speech. However, the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Rui Wang , Liping Chen , Kong Aik Lee , Zhengpeng Zha , Zhenhua Ling

Automatic Speaker Recognition Systems (SRSs) have been widely used in voice applications for personal identification and access control. A typical SRS consists of three stages, i.e., training, enrollment, and recognition. Previous work has…

Sound · Computer Science 2023-12-12 Xinfeng Li , Junning Ze , Chen Yan , Yushi Cheng , Xiaoyu Ji , Wenyuan Xu

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-20 Sefik Emre Eskimez , Takuya Yoshioka , Huaming Wang , Xiaofei Wang , Zhuo Chen , Xuedong Huang

Deep speech classification has achieved tremendous success and greatly promoted the emergence of many real-world applications. However, backdoor attacks present a new security threat to it, particularly with untrustworthy third-party…

Sound · Computer Science 2023-08-15 Zhe Ye , Terui Mao , Li Dong , Diqun Yan
‹ Prev 1 8 9 10 Next ›