English
Related papers

Related papers: CallShield: Secure Caller Authentication over Real…

200 papers

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Yang Yang , Yury Kartynnik , Yunpeng Li , Jiuqiang Tang , Xing Li , George Sung , Matthias Grundmann

As users increasingly rely on cloud-based computing services, it is important to ensure that uploaded speech data remains private. Existing solutions rely either on server-side methods or focus on hiding speaker identity. While these…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-26 Peter Wu , Paul Pu Liang , Jiatong Shi , Ruslan Salakhutdinov , Shinji Watanabe , Louis-Philippe Morency

In this work, we introduce the first autoregressive framework for real-time, audio-driven portrait animation, a.k.a, talking head. Beyond the challenge of lengthy animation times, a critical challenge in realistic talking head generation…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Dingcheng Zhen , Shunshun Yin , Shiyang Qin , Hou Yi , Ziwei Zhang , Siyuan Liu , Gan Qi , Ming Tao

With the rapid advancement of smart glasses, voice interaction has been widely adopted due to its naturalness and convenience. However, its practical deployment is often undermined by vulnerability to spoofing attacks, while no public…

Human-Computer Interaction · Computer Science 2026-05-12 Weiye Xu , Zhang Jiang , Siqi Zheng , Xiyuxing Zhang , Changhao Zhang , Jian Liu , Weiqiang Wang , Yuntao Wang

Recent results in end-to-end automatic speech recognition have demonstrated the efficacy of pseudo-labeling for semi-supervised models trained both with Connectionist Temporal Classification (CTC) and Sequence-to-Sequence (seq2seq) losses.…

Computation and Language · Computer Science 2021-08-31 Tatiana Likhomanenko , Qiantong Xu , Jacob Kahn , Gabriel Synnaeve , Ronan Collobert

Large language models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks, demonstrating human-level performance in text generation, reasoning, and question answering. However, training such…

Cryptography and Security · Computer Science 2025-11-17 Yanbo Dai , Zongjie Li , Zhenlan Ji , Shuai Wang

Audio and speech data are increasingly used in machine learning applications such as speech recognition, speaker identification, and mental health monitoring. However, the passive collection of this data by audio listening devices raises…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-16 Tu Duyen Nguyen , Adrien Lesage , Clotilde Cantini , Rachid Riad

Prevailing user authentication schemes on smartphones rely on explicit user interaction, where a user types in a passcode or presents a biometric cue such as face, fingerprint, or iris. In addition to being cumbersome and obtrusive to the…

Computer Vision and Pattern Recognition · Computer Science 2019-01-17 Debayan Deb , Arun Ross , Anil K. Jain , Kwaku Prakah-Asante , K. Venkatesh Prasad

Morphing techniques generate artificial biometric samples that combine features from multiple individuals, allowing each contributor to be verified against a single enrolled template. While extensively studied in face recognition, this…

Sound · Computer Science 2026-01-30 Bharath Krishnamurthy , Ajita Rattani

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by…

Sound · Computer Science 2023-10-18 Fernando López , Jordi Luque , Carlos Segura , Pablo Gómez

We describe a large vocabulary speech recognition system that is accurate, has low latency, and yet has a small enough memory and computational footprint to run faster than real-time on a Nexus 5 Android smartphone. We employ a quantized…

This paper describes a new baseline system for automatic speech recognition (ASR) in the CHiME-4 challenge to promote the development of noisy ASR in speech processing communities by providing 1) state-of-the-art system with a simplified…

Sound · Computer Science 2018-03-28 Szu-Jui Chen , Aswin Shanmugam Subramanian , Hainan Xu , Shinji Watanabe

In today's world, computer networks have become vulnerable to numerous attacks. In both wireless and wired networks, one of the most common attacks is man-in-the-middle attacks, within which session hijacking, context confusion attacks have…

Cryptography and Security · Computer Science 2022-02-02 Kailash Gogineni , Yongsheng Mei , Guru Venkataramani , Tian Lan

The adoption of Large Language Models (LLMs) has revolutionized AI applications but poses significant challenges in safeguarding user privacy. Ensuring compliance with privacy regulations such as GDPR and CCPA while addressing nuanced…

Cryptography and Security · Computer Science 2025-01-23 Shubhi Asthana , Bing Zhang , Ruchi Mahindru , Chad DeLuca , Anna Lisa Gentile , Sandeep Gopisetty

The widespread deployment of LLMs across enterprise services has created a critical security blind spot. Organizations operate multiple LLM services handling billions of queries daily, yet regulatory compliance boundaries prevent these…

Cryptography and Security · Computer Science 2026-03-03 Waris Gill , Natalie Isak , Matthew Dressman

We introduce a low-latency telecom AI voice agent pipeline for real-time, interactive telecommunications use, enabling advanced voice AI for call center automation, intelligent IVR (Interactive Voice Response), and AI-driven customer…

Sound · Computer Science 2025-08-08 Vignesh Ethiraj , Ashwath David , Sidhanth Menon , Divya Vijay

Securing the Internet of Things (IoT) is a necessary milestone toward expediting the deployment of its applications and services. In particular, the functionality of the IoT devices is extremely dependent on the reliability of their message…

Information Theory · Computer Science 2017-11-07 Aidin Ferdowsi , Walid Saad

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major…

It is well known that speaker identification yields very high performance in a neutral talking environment, on the other hand, the performance has been sharply declined in a shouted talking environment. This work aims at proposing,…

Sound · Computer Science 2017-07-07 Ismail Shahin

Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speaker feature…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Nikita Kuzmin , Songting Liu , Kong Aik Lee , Eng Siong Chng