English
Related papers

Related papers: CallShield: Secure Caller Authentication over Real…

200 papers

Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal restraint and enabling discrete interaction in noise-sensitive…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Cancan Li , Fei Su , Juan Liu , Hui Bu , Yulong Wan , Hongbin Suo , Ming Li

Earphones can give access to sensitive information via voice assistants which demands security methods that prevent unauthorized use. Therefore, we developed EarCapAuth, an authentication mechanism using 48 capacitive electrodes embedded…

Cryptography and Security · Computer Science 2024-11-08 Richard Hanser , Tobias Röddiger , Till Riedel , Michael Beigl

Deepfake speech attribution remains challenging for existing solutions. Classifier-based solutions often fail to generalize to domain-shifted samples, and watermarking-based solutions are easily compromised by distortions like codec…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-16 Wanying Ge , Xin Wang , Junichi Yamagishi

Large Language Models (LLMs) can now solve entire exams directly from uploaded PDF assessments, raising urgent concerns about academic integrity and the reliability of grades and credentials. Existing watermarking techniques either operate…

Computation and Language · Computer Science 2026-01-19 Ashish Raj Shekhar , Shiven Agarwal , Priyanuj Bordoloi , Yash Shah , Tejas Anvekar , Vivek Gupta

We use the term re-identification to refer to the process of recovering the original speaker's identity from anonymized speech outputs. Speaker de-identification systems aim to reduce the risk of re-identification, but most evaluations…

Sound · Computer Science 2025-09-19 Seungmin Seo , Oleg Aulov , P. Jonathon Phillips

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-23 Wenxi Chen , Ziyang Ma , Ruiqi Yan , Yuzhe Liang , Xiquan Li , Ruiyang Xu , Zhikang Niu , Yanqiao Zhu , Yifan Yang , Zhanxun Liu , Kai Yu , Yuxuan Hu , Jinyu Li , Yan Lu , Shujie Liu , Xie Chen

In recent years, the need for privacy preservation when manipulating or storing personal data, including speech , has become a major issue. In this paper, we present a system addressing the speaker-level anonymization problem. We propose…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-29 Francesco Nespoli , Daniel Barreda , Joerg Bitzer , Patrick A. Naylor

In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of different languages together for multilingual generalization…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Roza Chojnacka , Jason Pelecanos , Quan Wang , Ignacio Lopez Moreno

Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Hui-Peng Du , Yang Ai , Xiao-Hang Jiang , Yuan Tian , Zhen-Hua Ling

Speaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments. However, existing approaches extract utterance-level speaker embeddings from each channel of an…

Sound · Computer Science 2022-03-29 Chengdong Liang , Yijiang Chen , Jiadi Yao , Xiao-Lei Zhang

Invisible watermarking is essential for tracing the provenance of digital content. However, training state-of-the-art models remains notoriously difficult, with current approaches often struggling to balance robustness against true…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Tomáš Souček , Pierre Fernandez , Hady Elsahar , Sylvestre-Alvise Rebuffi , Valeriu Lacatusu , Tuan Tran , Tom Sander , Alexandre Mourachko

Speaker anonymization aims to protect the privacy of speakers while preserving spoken linguistic information from speech. Current mainstream neural network speaker anonymization systems are complicated, containing an F0 extractor, speaker…

Sound · Computer Science 2022-04-28 Xiaoxiao Miao , Xin Wang , Erica Cooper , Junichi Yamagishi , Natalia Tomashenko

Deep learning technology has made it possible to generate realistic content of specific individuals. These `deepfakes' can now be generated in real-time which enables attackers to impersonate people over audio and video calls. Moreover,…

Cryptography and Security · Computer Science 2023-01-10 Lior Yasur , Guy Frankovits , Fred M. Grabovski , Yisroel Mirsky

The popularity of mobile device has made people's lives more convenient, but threatened people's privacy at the same time. As end users are becoming more and more concerned on the protection of their private information, it is even harder…

Cryptography and Security · Computer Science 2014-07-04 Zhe Zhou , Wenrui Diao , Xiangyu Liu , Kehuan Zhang

This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings. By…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-16 Weiqing Wang , Ming Li

Cache attacks pose a threat to any code whose execution flow or memory accesses depend on sensitive information. Especially in public clouds, where caches are shared across several tenants, cache attacks remain an unsolved problem. Cache…

Cryptography and Security · Computer Science 2017-09-07 Samira Briongos , Gorka Irazoqui , Pedro Malagón , Thomas Eisenbarth

Recent advances in AudioLLMs have enabled spoken dialogue systems to move beyond turn-based interaction toward real-time full-duplex communication, where the agent must decide when to speak, yield, or interrupt while the user is still…

Artificial Intelligence Generated Content (AIGC) has advanced significantly, particularly with the development of video generation models such as text-to-video (T2V) models and image-to-video (I2V) models. However, like other AIGC types,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Runyi Hu , Jie Zhang , Yiming Li , Jiwei Li , Qing Guo , Han Qiu , Tianwei Zhang

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved…

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain
‹ Prev 1 8 9 10 Next ›