English
Related papers

Related papers: A Study On Convolutional Neural Network Based End-…

200 papers

A great deal of recent research effort on speech spoofing countermeasures has been invested into back-end neural networks and training criteria. We contribute to this effort with a comparative perspective in this study. Our comparison of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Xin Wang , Junich Yamagishi

We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous…

Computation and Language · Computer Science 2016-06-21 Liang Lu , Lingpeng Kong , Chris Dyer , Noah A. Smith , Steve Renals

Generalized end-to-end (GE2E) model is widely used in speaker verification (SV) fields due to its expandability and generality regardless of specific languages. However, the long-short term memory (LSTM) based on GE2E has two limitations:…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Hyeonmook Park , Jungbae Park , Sang Wan Lee

It is now well-known that automatic speaker verification (ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofing countermeasure…

Cryptography and Security · Computer Science 2024-01-30 Xuechen Liu , Md Sahidullah , Kong Aik Lee , Tomi Kinnunen

In this paper, we evaluate convolutional neural network (CNN) features using the AlexNet architecture and very deep convolutional network (VGGNet) architecture. To date, most CNN researchers have employed the last layers before output,…

Computer Vision and Pattern Recognition · Computer Science 2015-09-28 Hirokatsu Kataoka , Kenji Iwata , Yutaka Satoh

In speaker verification, the extraction of voice representations is mainly based on the Residual Neural Network (ResNet) architecture. ResNet is built upon convolution layers which learn filters to capture local spatial patterns along all…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-14 Mickael Rouvier , Pierre-Michel Bousquet

The wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-05 Yuxiang Zhang , Jingze Lu , Zengqiang Shang , Wenchao Wang , Pengyuan Zhang

This paper presents the Speech Technology Center (STC) systems submitted to Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) Challenge 2015. In this work we investigate different acoustic feature spaces to determine…

The ability of countermeasure models to generalize from seen speech synthesis methods to unseen ones has been investigated in the ASVspoof challenge. However, a new mismatch scenario in which fake audio may be generated from real audio with…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Chang Zeng , Xin Wang , Xiaoxiao Miao , Erica Cooper , Junichi Yamagishi

Deep learning is progressively gaining popularity as a viable alternative to i-vectors for speaker recognition. Promising results have been recently obtained with Convolutional Neural Networks (CNNs) when fed by raw speech samples directly.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-12 Mirco Ravanelli , Yoshua Bengio

Convolutional neural networks (CNNs) are similar to "ordinary" neural networks in the sense that they are made up of hidden layers consisting of neurons with "learnable" parameters. These neurons receive inputs, performs a dot product, and…

Computer Vision and Pattern Recognition · Computer Science 2019-02-08 Abien Fred Agarap

In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Jose A. Lopez , Georg Stemmer , Héctor Cordourier Maruri

The ASVspoof initiative was conceived to spearhead research in anti-spoofing for automatic speaker verification (ASV). This paper describes the third in a series of bi-annual challenges: ASVspoof 2019. With the challenge database and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Andreas Nautsch , Xin Wang , Nicholas Evans , Tomi Kinnunen , Ville Vestman , Massimiliano Todisco , Héctor Delgado , Md Sahidullah , Junichi Yamagishi , Kong Aik Lee

Understanding accent is an issue which can derail any human-machine interaction. Accent classification makes this task easier by identifying the accent being spoken by a person so that the correct words being spoken can be identified by…

Sound · Computer Science 2019-10-16 Asad Ahmed , Pratham Tangri , Anirban Panda , Dhruv Ramani , Samarjit Karmakar

This paper describes the NPU system submitted to Spoofing Aware Speaker Verification Challenge 2022. We particularly focus on the \textit{backend ensemble} for speaker verification and spoofing countermeasure from three aspects. Firstly,…

Sound · Computer Science 2022-09-26 Li Zhang , Yue Li , Huan Zhao , Qing Wang , Lei Xie

Motivated by the fact that characteristics of different sound classes are highly diverse in different temporal scales and hierarchical levels, a novel deep convolutional neural network (CNN) architecture is proposed for the environmental…

Sound · Computer Science 2018-06-15 Boqing Zhu , Kele Xu , Dezhi Wang , Lilun Zhang , Bo Li , Yuxing Peng

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker…

Sound · Computer Science 2024-12-25 Xuechen Liu , Junichi Yamagishi , Md Sahidullah , Tomi kinnunen

Recent years have witnessed the great success of convolutional neural network (CNN) based models in the field of computer vision. CNN is able to learn hierarchically abstracted features from images in an end-to-end training manner. However,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-16 Xin Li , Zequn Jie , Jiashi Feng , Changsong Liu , Shuicheng Yan

We propose a novel method for Acoustic Event Detection (AED). In contrast to speech, sounds coming from acoustic events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an extended time…

Sound · Computer Science 2016-12-09 Naoya Takahashi , Michael Gygli , Beat Pfister , Luc Van Gool

Spoofing-aware speaker verification (SASV) jointly addresses automatic speaker verification and spoofing countermeasures to improve robustness against adversarial attacks. In this paper, we investigate our recently proposed modular SASV…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-03 Oguzhan Kurnaz , Jagabandhu Mishra , Tomi Kinnunen , Cemal Hanilci
‹ Prev 1 4 5 6 7 8 10 Next ›