English
Related papers

Related papers: Replay attack detection with complementary high-re…

200 papers

Automatic speaker verification (ASV) systems utilize the biometric information in human speech to verify the speaker's identity. The techniques used for performing speaker verification are often vulnerable to malicious attacks that attempt…

Sound · Computer Science 2020-11-26 Yang Gao , Jiachen Lian , Bhiksha Raj , Rita Singh

In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech…

Sound · Computer Science 2023-05-31 Haoyu Lu , Nan Li , Tongtong Song , Longbiao Wang , Jianwu Dang , Xiaobao Wang , Shiliang Zhang

Deep neural network (DNN)-based joint source and channel coding is proposed for privacy-aware end-to-end image transmission against multiple eavesdroppers. Both scenarios of colluding and non-colluding eavesdroppers are considered. Unlike…

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Component-level audio Spoofing (Comp-Spoof) targets a new form of audio manipulation where only specific components of a signal, such as speech or environmental sound, are forged or substituted while other components remain genuine.…

Sound · Computer Science 2026-02-02 Xueping Zhang , Yechen Wang , Linxi Li , Liwei Jin , Ming Li

Distant speech recognition is a challenge, particularly due to the corruption of speech signals by reverberation caused by large distances between the speaker and microphone. In order to cope with a wide range of reverberations in…

Computation and Language · Computer Science 2016-08-18 Jeehye Lee , Myungin Lee , Joon-Hyuk Chang

This paper proposes a deep speech enhancement method which exploits the high potential of residual connections in a wide neural network architecture, a topology known as Wide Residual Network. This is supported on single dimensional…

Sound · Computer Science 2019-01-04 Dayana Ribas , Jorge Llombart , Antonio Miguel , Luis Vicente

The existing fake audio detection systems often rely on expert experience to design the acoustic features or manually design the hyperparameters of the network structure. However, artificial adjustment of the parameters can have a…

The development of deep learning technology has greatly promoted the performance improvement of automatic speech recognition (ASR) technology, which has demonstrated an ability comparable to human hearing in many tasks. Voice interfaces are…

Sound · Computer Science 2022-06-09 Jinghui Xu , Jifeng Zhu , Yong Yang

Voice interfaces are becoming accepted widely as input methods for a diverse set of devices. This development is driven by rapid improvements in automatic speech recognition (ASR), which now performs on par with human listening in many…

Cryptography and Security · Computer Science 2018-10-31 Lea Schönherr , Katharina Kohls , Steffen Zeiler , Thorsten Holz , Dorothea Kolossa

We propose a stacked 1D convolutional neural network (S1DCNN) for end-to-end small footprint voice trigger detection in a streaming scenario. Voice trigger detection is an important speech application, with which users can activate their…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Takuya Higuchi , Mohammad Ghasemzadeh , Kisun You , Chandra Dhir

Replay attacks belong to the class of severe threats against voice-controlled systems, exploiting the easy accessibility of speech signals by recorded and replayed speech to grant unauthorized access to sensitive data. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-09 Michael Neri , Tuomas Virtanen

Deep neural network (DNN) based speech enhancement models have attracted extensive attention due to their promising performance. However, it is difficult to deploy a powerful DNN in real-time applications because of its high computational…

Sound · Computer Science 2022-07-25 Xiaohuai Le , Tong Lei , Kai Chen , Jing Lu

Many endeavors have sought to develop countermeasure techniques as enhancements on Automatic Speaker Verification (ASV) systems, in order to make them more robust against spoof attacks. As evidenced by the latest ASVspoof 2019…

Sound · Computer Science 2021-09-21 Amir Mohammad Rostami , Mohammad Mehdi Homayounpour , Ahmad Nickabadi

The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of track 1 (Low-quality Fake Audio Detection) and track 2…

Sound · Computer Science 2022-10-12 Xiaohui Liu , Meng Liu , Lin Zhang , Linjuan Zhang , Chang Zeng , Kai Li , Nan Li , Kong Aik Lee , Longbiao Wang , Jianwu Dang

Employing deep neural networks (DNNs) to directly learn filters for multi-channel speech enhancement has potentially two key advantages over a traditional approach combining a linear spatial filter with an independent tempo-spectral…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-23 Kristina Tesch , Nils-Hendrik Mohrmann , Timo Gerkmann

In this paper, we present UR-AIR system submission to the logical access (LA) and the speech deepfake (DF) tracks of the ASVspoof 2021 Challenge. The LA and DF tasks focus on synthetic speech detection (SSD), i.e. detecting text-to-speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-05 Xinhui Chen , You Zhang , Ge Zhu , Zhiyao Duan

End-to-end speech recognition systems have achieved competitive results compared to traditional systems. However, the complex transformations involved between layers given highly variable acoustic signals are hard to analyze. In this paper,…

Computation and Language · Computer Science 2019-11-05 Chung-Yi Li , Pei-Chieh Yuan , Hung-Yi Lee

The field of speech separation, addressing the "cocktail party problem", has seen revolutionary advances with DNNs. Speech separation enhances clarity in complex acoustic environments and serves as crucial pre-processing for speech…

Sound · Computer Science 2025-08-15 Kai Li , Guo Chen , Wendi Sang , Yi Luo , Zhuo Chen , Shuai Wang , Shulin He , Zhong-Qiu Wang , Andong Li , Zhiyong Wu , Xiaolin Hu

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker…

Sound · Computer Science 2024-12-25 Xuechen Liu , Junichi Yamagishi , Md Sahidullah , Tomi kinnunen