中文
相关论文

相关论文: CallShield: Secure Caller Authentication over Real…

200 篇论文

With the rapid advancement of speech generative models, unauthorized voice cloning poses significant privacy and security risks. Speech watermarking offers a viable solution for tracing sources and preventing misuse. Current watermarking…

声音 · 计算机科学 2025-09-30 Yang Cui , Peter Pan , Lei He , Sheng Zhao

Facilitated by the speech generation framework that disentangles speech into content, speaker, and prosody, voice anonymization is accomplished by substituting the original speaker embedding vector with that of a pseudo-speaker. In this…

音频与语音处理 · 电气工程与系统科学 2025-12-04 Zeyan Liu , Liping Chen , Kong Aik Lee , Zhenhua Ling

Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages. However, their application to long audio transcription via buffered or…

声音 · 计算机科学 2023-07-12 Max Bain , Jaesung Huh , Tengda Han , Andrew Zisserman

Sound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based SSL methods provide analytic solutions under specific signal and noise assumptions,…

声音 · 计算机科学 2024-09-12 Xinyuan Qian , Xianghu Yue , Jiadong Wang , Huiping Zhuang , Haizhou Li

Speaker recognition is a biometric modality that uses underlying speech information to determine the identity of the speaker. Speaker Identification (SID) under noisy conditions is one of the challenging topics in the field of speech…

声音 · 计算机科学 2019-08-02 Nursadul Mamun , Ria Ghosh , John H. L. Hansen

Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio offers spatial cues that can enhance replay detection…

音频与语音处理 · 电气工程与系统科学 2025-09-19 Michael Neri , Tuomas Virtanen

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Ivan Kukanov , Jun Wah Ng

We propose SpeakerNet - a new neural architecture for speaker recognition and speaker verification tasks. It is composed of residual blocks with 1D depth-wise separable convolutions, batch-normalization, and ReLU layers. This architecture…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Nithin Rao Koluguri , Jason Li , Vitaly Lavrukhin , Boris Ginsburg

Caller ID spoofing is a global industry problem and often acts as a critical enabler for telephone fraud. To address this problem, the Federal Communications Commission (FCC) has mandated telecom providers in the US to implement…

密码学与安全 · 计算机科学 2023-09-26 Shen Wang , Mahshid Delavar , Muhammad Ajmal Azad , Farshad Nabizadeh , Steve Smith , Feng Hao

Voice user interfaces (VUIs) are rapidly transitioning from accessibility features to mainstream interaction modalities. Yet most operating systems' built-in voice commands remain underutilized despite possessing robust technical…

人机交互 · 计算机科学 2026-02-27 Md Ehtesham-Ul-Haque , Syed Masum Billah

Real-time automatic speech recognition (ASR) systems face a fundamental trade-off between transcription accuracy and computational efficiency, particularly when deploying large-scale transformer models like Whisper. Existing streaming…

This paper presents the design and implementation of WhisperWand, a comprehensive voice and motion tracking interface for voice assistants. Distinct from prior works, WhisperWand is a precise tracking interface that can co-exist with the…

人机交互 · 计算机科学 2023-01-26 Yang Bai , Irtaza Shahid , Harshvardhan Takawale , Nirupam Roy

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

声音 · 计算机科学 2024-02-08 Dan Lyth , Simon King

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

The emergence of voice-assistant devices ushers in delightful user experiences not just on the smart home front, but also in diverse educational environments from classrooms to personalized-learning/tutoring. However, the use of voice as an…

音频与语音处理 · 电气工程与系统科学 2021-04-23 Mohammad Niknazar , Aditya Vempaty , Ravi Kokku

Deepfakes and manipulated media are becoming a prominent threat due to the recent advances in realistic image and video synthesis techniques. There have been several attempts at combating Deepfakes using machine learning classifiers.…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Paarth Neekhara , Shehzeen Hussain , Xinqiao Zhang , Ke Huang , Julian McAuley , Farinaz Koushanfar

Automatic speech recognition and voice identification systems are being deployed in a wide array of applications, from providing control mechanisms to devices lacking traditional interfaces, to the automatic transcription of conversations…

The presence of non-speech segments in utterances often leads to the performance degradation of speaker verification. Existing systems usually use voice activation detection as a preprocessing step to cut off long silence segments. However,…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Zijun Huang , Chengdong Liang , Jiadi Yao , Xiao-Lei Zhang

The proliferation of AI-generated content and sophisticated video editing tools has made it both important and challenging to moderate digital platforms. Video watermarking addresses these challenges by embedding imperceptible signals into…

多媒体 · 计算机科学 2024-12-13 Pierre Fernandez , Hady Elsahar , I. Zeki Yalniz , Alexandre Mourachko

In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is to embed signatures that can be detected with high accuracy…

声音 · 计算机科学 2025-04-16 Patrick O'Reilly , Zeyu Jin , Jiaqi Su , Bryan Pardo