中文
相关论文

相关论文: Evaluating the Representation of Vowels in Wav2Vec…

200 篇论文

This paper proposes a novel Wavelet Packet based feature extraction approach for the task of text independent speaker recognition. The features are extracted by using the combination of Mel Frequency Cepstral Coefficient (MFCC) and Wavelet…

In this article, we conduct a study on the performance of some supervised learning algorithms for vowel recognition. This study aims to compare the accuracy of each algorithm. Thus, we present an empirical comparison between five supervised…

计算与语言 · 计算机科学 2015-07-23 Rimah Amami , Dorra Ben Ayed , Noureddine Ellouze

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand…

音频与语音处理 · 电气工程与系统科学 2018-09-13 Suwon Shon , Hao Tang , James Glass

This paper tests the hypothesis that distinctive feature classifiers anchored at phonetic landmarks can be transferred cross-lingually without loss of accuracy. Three consonant voicing classifiers were developed: (1) manually selected…

计算与语言 · 计算机科学 2017-08-23 Xiang Kong , Xuesong Yang , Mark Hasegawa-Johnson , Jeung-Yoon Choi , Stefanie Shattuck-Hufnagel

With the success of neural network based modeling in automatic speech recognition (ASR), many studies investigated acoustic modeling and learning of feature extractors directly based on the raw waveform. Recently, one line of research has…

音频与语音处理 · 电气工程与系统科学 2021-10-06 Peter Vieting , Christoph Lüscher , Wilfried Michel , Ralf Schlüter , Hermann Ney

Owing to large-scale image-text contrastive training, pre-trained vision language model (VLM) like CLIP shows superior open-vocabulary recognition ability. Most existing open-vocabulary object detectors attempt to utilize the pre-trained…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Xiangyu Gao , Yu Dai , Benliu Qiu , Lanxiao Wang , Heqian Qiu , Hongliang Li

Deep neural networks are representation learning techniques. During training, a deep net is capable of generating a descriptive language of unprecedented size and detail in machine learning. Extracting the descriptive language coded within…

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extractors for a broader…

音频与语音处理 · 电气工程与系统科学 2025-11-21 Wei-Cheng Tseng , David Harwath

Neural network-based vocoders have recently demonstrated the powerful ability to synthesize high-quality speech. These models usually generate samples by conditioning on spectral features, such as Mel-spectrogram and fundamental frequency,…

音频与语音处理 · 电气工程与系统科学 2023-03-13 Yunchao He , Yujun Wang

Speech Emotion Recognition is a crucial area of research in human-computer interaction. While significant work has been done in this field, many state-of-the-art networks struggle to accurately recognize emotions in speech when the data is…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Rashedul Hasan , Meher Nigar , Nursadul Mamun , Sayan Paul

Speech is the most natural way of expressing ourselves as humans. Identifying emotion from speech is a nontrivial task due to the ambiguous definition of emotion itself. Speaker Emotion Recognition (SER) is essential for understanding human…

声音 · 计算机科学 2024-11-07 Pourya Jafarzadeh , Amir Mohammad Rostami , Padideh Choobdar

Deep learning is still not a very common tool in speaker verification field. We study deep convolutional neural network performance in the text-prompted speaker verification task. The prompted passphrase is segmented into word states - i.e.…

音频与语音处理 · 电气工程与系统科学 2018-03-15 Sergey Novoselov , Oleg Kudashev , Vadim Schemelinin , Ivan Kremnev , Galina Lavrentyeva

This project intends to study the image representation based on attention mechanism and multimodal data. By adding multiple pattern layers to the attribute model, the semantic and hidden layers of image content are integrated. The word…

计算与语言 · 计算机科学 2024-06-14 Dan Sun , Yaxin Liang , Yining Yang , Yuhan Ma , Qishi Zhan , Erdi Gao

In recent years, using raw waveforms as input for deep networks has been widely explored for the speaker verification system. For example, RawNet and RawNet2 extracted speaker's feature embeddings from waveforms automatically for…

音频与语音处理 · 电气工程与系统科学 2021-10-08 Jin Li , Nan Yan , Lan Wang

In recent years, deep learning-based models have significantly improved the Natural Language Processing (NLP) tasks. Specifically, the Convolutional Neural Network (CNN), initially used for computer vision, has shown remarkable performance…

计算与语言 · 计算机科学 2022-03-11 Sanskar Soni , Satyendra Singh Chouhan , Santosh Singh Rathore

Automatic speaker verification (ASV) systems are often affected by spoofing attacks. Recent transformer-based models have improved anti-spoofing performance by learning strong feature representations. However, these models usually need high…

音频与语音处理 · 电气工程与系统科学 2025-07-14 Yang Xiao , Ting Dang , Rohan Kumar Das

Vocal tract configurations play a vital role in generating distinguishable speech sounds, by modulating the airflow and creating different resonant cavities in speech production. They contain abundant information that can be utilized to…

声音 · 计算机科学 2018-07-31 Pramit Saha , Praneeth Srungarapu , Sidney Fels

We propose a learnable mel-frequency cepstral coefficient (MFCC) frontend architecture for deep neural network (DNN) based automatic speaker verification. Our architecture retains the simplicity and interpretability of MFCC-based features…

声音 · 计算机科学 2021-02-23 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Automatic speech recognition (ASR) systems typically use handcrafted feature extraction pipelines. To avoid their inherent information loss and to achieve more consistent modeling from speech to transcribed text, neural raw waveform feature…

音频与语音处理 · 电气工程与系统科学 2023-08-09 Peter Vieting , Ralf Schlüter , Hermann Ney

Interpretability work on the convolutional layers of CNNs has primarily focused on computer vision, but some studies also explore correspondences between the latent space and the output in the audio domain. However, it has not been…

计算与语言 · 计算机科学 2025-01-15 Bruno Ferenc Šegedin , Gasper Beguš