中文
相关论文

相关论文: A Practical Guide to Logical Access Voice Presenta…

200 篇论文

Current state-of-the-art automatic speaker verification (ASV) systems are vulnerable to presentation attacks, and several countermeasures (CMs), which distinguish bona fide trials from spoofing ones, have been explored to protect ASV.…

音频与语音处理 · 电气工程与系统科学 2022-10-27 Chang Zeng , Lin Zhang , Meng Liu , Junichi Yamagishi

\textit{Objective:} Conventional EEG-based auditory attention detection (AAD) is achieved by comparing the time-varying speech stimuli and the elicited EEG signals. However, in order to obtain reliable correlation values, these methods…

音频与语音处理 · 电气工程与系统科学 2023-08-30 Hongxu Zhu , Siqi Cai , Yidi Jiang , Qiquan Zhang , Haizhou Li

In recent years, the popularity of fingerprint-based biometric authentication systems significantly increased. However, together with many advantages, biometric systems are still vulnerable to presentation attacks (PAs). In particular, this…

计算机视觉与模式识别 · 计算机科学 2021-01-25 Jascha Kolberg , Marcel Grimmer , Marta Gomez-Barrero , Christoph Busch

The objective of automatic speaker verification (ASV) systems is to determine whether a given test speech utterance corresponds to a claimed enrolled speaker. These systems have a wide range of applications, and ensuring their reliability…

音频与语音处理 · 电气工程与系统科学 2025-05-27 Amro Asali , Yehuda Ben-Shimol , Itshak Lapidot

Current Active Speaker Detection (ASD) models achieve great results on AVA-ActiveSpeaker (AVA), using only sound and facial features. Although this approach is applicable in movie setups (AVA), it is not suited for less constrained…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

This paper proposes a method for utilizing thermal features of the hand for the purpose of presentation attack detection (PAD) that can be employed in a hand biometrics system's pipeline. By envisaging two different operational modes of our…

计算机视觉与模式识别 · 计算机科学 2018-09-13 Ewelina Bartuzi , Mateusz Trokielewicz

Nowadays, the development of a Presentation Attack Detection (PAD) system for ID cards presents a challenge due to the lack of images available to train a robust PAD system and the increase in diversity of possible attack instrument…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Qingwen Zeng , Juan E. Tapia , Izan Garcia , Juan M. Espin , Christoph Busch

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial in applications such…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Junhua Liao , Haihan Duan , Kanghui Feng , Wanbing Zhao , Yanbing Yang , Liangyin Chen

As a versatile AI application, voice assistants (VAs) have become increasingly popular, but are vulnerable to security threats. Attackers have proposed various inaudible attacks, but are limited by cost, distance, or LoS. Therefore, we…

密码学与安全 · 计算机科学 2026-03-26 Chao Liu , Zhezheng Zhu , Hao Chen , Kaiwen Guo , Penghao Wang , Xiang-Yang Li

Recent advances in text-to-speech (TTS) systems, particularly those with voice cloning capabilities, have made voice impersonation readily accessible, raising ethical and legal concerns due to potential misuse for malicious activities like…

声音 · 计算机科学 2024-10-10 Hongbin Liu , Youzheng Chen , Arun Narayanan , Athula Balachandran , Pedro J. Moreno , Lun Wang

The diffusion of fingerprint verification systems for security applications makes it urgent to investigate the embedding of software-based presentation attack detection algorithms (PAD) into such systems. Companies and institutions need to…

密码学与安全 · 计算机科学 2021-10-22 Marco Micheletto , Gian Luca Marcialis , Giulia Orrù , Fabio Roli

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

声音 · 计算机科学 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan

Foundation models are becoming increasingly popular due to their strong generalization capabilities resulting from being trained on huge datasets. These generalization capabilities are attractive in areas such as NIR Iris Presentation…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Juan E. Tapia , Lázaro Janier González-Soler , Christoph Busch

Biometric systems are vulnerable to Presentation Attacks (PA) performed using various Presentation Attack Instruments (PAIs). Even though there are numerous Presentation Attack Detection (PAD) techniques based on both deep learning and…

计算机视觉与模式识别 · 计算机科学 2023-06-05 Zhe Kong , Wentian Zhang , Feng Liu , Wenhan Luo , Haozhe Liu , Linlin Shen , Raghavendra Ramachandra

In speech technologies, speaker's voice representation is used in many applications such as speech recognition, voice conversion, speech synthesis and, obviously, user authentication. Modern vocal representations of the speaker are based on…

音频与语音处理 · 电气工程与系统科学 2021-06-17 Paul-Gauthier Noé , Mohammad Mohammadamini , Driss Matrouf , Titouan Parcollet , Andreas Nautsch , Jean-François Bonastre

High-performance anti-spoofing models for automatic speaker verification (ASV), have been widely used to protect ASV by identifying and filtering spoofing audio that is deliberately generated by text-to-speech, voice conversion, audio…

音频与语音处理 · 电气工程与系统科学 2020-12-08 Haibin Wu , Andy T. Liu , Hung-yi Lee

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

计算与语言 · 计算机科学 2023-05-15 Fei Tao , Carlos Busso

Background noise reduces speech intelligibility and quality, making speaker verification (SV) in noisy environments a challenging task. To improve the noise robustness of SV systems, additive noise data augmentation method has been commonly…

音频与语音处理 · 电气工程与系统科学 2023-07-21 Wonbin Kim , Hyun-seo Shin , Ju-ho Kim , Jungwoo Heo , Chan-yeong Lim , Ha-Jin Yu

Recent text-to-speech (TTS) developments have made voice cloning (VC) more realistic, affordable, and easily accessible. This has given rise to many potential abuses of this technology, including Joe Biden's New Hampshire deepfake robocall.…

音频与语音处理 · 电气工程与系统科学 2025-08-29 Hashim Ali , Surya Subramani , Hafiz Malik

Recent studies have demonstrated the vulnerability of Automatic Speech Recognition systems to adversarial examples, which can deceive these systems into misinterpreting input speech commands. While previous research has primarily focused on…

声音 · 计算机科学 2025-11-21 Aravindhan G , Yuvaraj Govindarajulu , Parin Shah