中文
相关论文

相关论文: Privacy-Preserving End-to-End Full-Duplex Speech D…

200 篇论文

We present an end-to-end speech recognition model that learns interaction between two speakers based on the turn-changing information. Unlike conventional speech recognition models, our model exploits two speakers' history of…

音频与语音处理 · 电气工程与系统科学 2019-07-26 Suyoun Kim , Siddharth Dalmia , Florian Metze

Invoking large transmit antenna arrays, massive MIMO wiretap settings are capable of suppressing passive eavesdroppers via narrow beamforming towards legitimate terminals. This implies that secrecy is obtained almost for free in these…

信息论 · 计算机科学 2019-12-06 Ali Bereyhi , Saba Asaad , Ralf R. Müller , Rafael F. Schaefer , Georg Fischer , H. Vincent Poor

Collecting speech data is an important step in training speech recognition systems and other speech-based machine learning models. However, the issue of privacy protection is an increasing concern that must be addressed. The current study…

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair benchmark and…

音频与语音处理 · 电气工程与系统科学 2025-02-05 Jixun Yao , Nikita Kuzmin , Qing Wang , Pengcheng Guo , Ziqian Ning , Dake Guo , Kong Aik Lee , Eng-Siong Chng , Lei Xie

New Advances in machine learning have made Automated Speech Recognition (ASR) systems practical and more scalable. These systems, however, pose serious privacy threats as speech is a rich source of sensitive acoustic and textual…

密码学与安全 · 计算机科学 2020-07-06 Shimaa Ahmed , Amrita Roy Chowdhury , Kassem Fawaz , Parmesh Ramanathan

This paper presents a number of fundamental properties of full-duplex radio for secure wireless communication under some simple and practical conditions. In particular, we consider the fields of secrecy capacity of a wireless channel…

信息论 · 计算机科学 2017-11-29 Yingbo Hua , Qiping Zhu , Reza Sohrabi

The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example by disrupting acoustic continuity or reducing vocal…

音频与语音处理 · 电气工程与系统科学 2026-04-21 Yunchong Xiao , Yuxiang Zhao , Ziyang Ma , Shuai Wang , Kai Yu , Jiachun Liao , Xie Chen

Anonymization of voice seeks to conceal the identity of the speaker while maintaining the utility of speech data. However, residual speaker cues often persist, which pose privacy risks. We propose SegReConcat, a data augmentation method for…

声音 · 计算机科学 2025-08-27 Ridwan Arefeen , Xiaoxiao Miao , Rong Tong , Aik Beng Ng , Simon See

Recent advances of end-to-end models have outperformed conventional models through employing a two-pass model. The two-pass model provides better speed-quality trade-offs for on-device speech recognition, where a 1st-pass model generates…

音频与语音处理 · 电气工程与系统科学 2020-09-24 Wei Li , James Qin , Chung-Cheng Chiu , Ruoming Pang , Yanzhang He

Neuromorphic vision sensors offer low latency and high dynamic range, but their deployment in public spaces raises severe data protection concerns. Recent Event-to-Video (E2V) models can reconstruct high-fidelity intensity images from…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Adam T. Müller , Mihai Kocsis , Nicolaj C. Stache

Voice user interfaces and digital assistants are rapidly entering our lives and becoming singular touch points spanning our devices. These always-on services capture and transmit our audio data to powerful cloud services for further…

计算与语言 · 计算机科学 2022-10-03 Ranya Aloufi , Hamed Haddadi , David Boyle

Multilingual end-to-end(E2E) models have shown a great potential in the expansion of the language coverage in the realm of automatic speech recognition(ASR). In this paper, we aim to enhance the multilingual ASR performance in two ways,…

计算与语言 · 计算机科学 2021-10-18 Rimita Lahiri , Kenichi Kumatani , Eric Sun , Yao Qian

In this paper, we explore the encoding/pooling layer and loss function in the end-to-end speaker and language recognition system. First, a unified and interpretable end-to-end system for both speaker and language recognition is developed.…

音频与语音处理 · 电气工程与系统科学 2018-04-17 Weicheng Cai , Jinkun Chen , Ming Li

Automatic speech recognition (ASR) is a key technology in many services and applications. This typically requires user devices to send their speech data to the cloud for ASR decoding. As the speech signal carries a lot of information about…

计算与语言 · 计算机科学 2019-11-13 Brij Mohan Lal Srivastava , Aurélien Bellet , Marc Tommasi , Emmanuel Vincent

In this work, the critical role of noisy feedback in enhancing the secrecy capacity of the wiretap channel is established. Unlike previous works, where a noiseless public discussion channel is used for feedback, the feed-forward and…

信息论 · 计算机科学 2016-11-18 Lifeng Lai , Hesham El Gamal , H. Vincent Poor

Voice assistive technologies have given rise to far-reaching privacy and security concerns. In this paper we investigate whether modular automatic speech recognition (ASR) can improve privacy in voice assistive systems by combining…

计算与语言 · 计算机科学 2021-04-05 Ranya Aloufi , Hamed Haddadi , David Boyle

With the popularity of virtual assistants (e.g., Siri, Alexa), the use of speech recognition is now becoming more and more widespread.However, speech signals contain a lot of sensitive information, such as the speaker's identity, which…

音频与语音处理 · 电气工程与系统科学 2022-03-21 Pierre Champion , Denis Jouvet , Anthony Larcher

We introduce a novel method to improve the performance of the VoicePrivacy Challenge 2022 baseline B1 variants. Among the known deficiencies of x-vector-based anonymization systems is the insufficient disentangling of the input features. In…

音频与语音处理 · 电气工程与系统科学 2022-11-01 Ünal Ege Gaznepoglu , Anna Leschanowsky , Nils Peters

End-to-end (E2E) spoken language understanding (SLU) systems that generate a semantic parse from speech have become more promising recently. This approach uses a single model that utilizes audio and text representations from pre-trained…

计算与语言 · 计算机科学 2023-07-25 Suyoun Kim , Akshat Shrivastava , Duc Le , Ju Lin , Ozlem Kalinli , Michael L. Seltzer

Transformer-based end-to-end neural speaker diarization (EEND) models utilize the multi-head self-attention (SA) mechanism to enable accurate speaker label prediction in overlapped speech regions. In this study, to enhance the training…

音频与语音处理 · 电气工程与系统科学 2023-03-03 Ye-Rin Jeoung , Joon-Young Yang , Jeong-Hwan Choi , Joon-Hyuk Chang