English
Related papers

Related papers: Wavelet-Based Mel-Frequency Cepstral Coefficients …

200 papers

Speaker recognition performance has been greatly improved with the emergence of deep learning. Deep neural networks show the capacity to effectively deal with impacts of noise and reverberation, making them attractive to far-field speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Wenda Chen , Jonathan Huang , Tobias Bocklet

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

Sound · Computer Science 2022-06-20 Yikang Wang , Hiromitsu Nishizaki

Determination of mispronunciations and ensuring feedback to users are maintained by computer-assisted language learning (CALL) systems. In this work, we introduce an ensemble model that defines the mispronunciation of Arabic phonemes and…

Sound · Computer Science 2023-01-05 Sukru Selim Calik , Ayhan Kucukmanisa , Zeynep Hilal Kilimci

In this paper, we propose a deep learning (DL)-based parameter enhancement method for a mixed excitation linear prediction (MELP) speech codec in noisy communication environment. Unlike conventional speech enhancement modules that are…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-21 Min-Jae Hwang , Hong-Goo Kang

State-of-the-art methods for audio generation suffer from fingerprint artifacts and repeated inconsistencies across temporal and spectral domains. Such artifacts could be well captured by the frequency domain analysis over the spectrogram.…

Sound · Computer Science 2021-06-29 Yang Gao , Tyler Vuong , Mahsa Elyasi , Gaurav Bharaj , Rita Singh

In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker's…

Sound · Computer Science 2024-06-03 Jungil Kong , Junmo Lee , Jeongmin Kim , Beomjeong Kim , Jihoon Park , Dohee Kong , Changheon Lee , Sangjin Kim

The direct expansion of deep neural network (DNN) based wide-band speech enhancement (SE) to full-band processing faces the challenge of low frequency resolution in low frequency range, which would highly likely lead to deteriorated…

Sound · Computer Science 2022-06-28 Zhongshu Hou , Qinwen Hu , Kai Chen , Jing Lu

DeepFake Audio, unlike DeepFake images and videos, has been relatively less explored from detection perspective, and the solutions which exist for the synthetic speech classification either use complex networks or dont generalize to…

Sound · Computer Science 2022-10-24 Vardhan Dongre , Abhinav Thimma Reddy , Nikhitha Reddeddy

The Vocal Joystick Vowel Corpus, by Washington University, was used to study monophthongs pronounced by native English speakers. The objective of this study was to quantitatively measure the extent at which speech recognition methods can…

Computation and Language · Computer Science 2017-02-24 Keith Y. Patarroyo , Vladimir Vargas-Calderón

Traditional automatic speech recognition (ASR) systems often use an acoustic model (AM) built on handcrafted acoustic features, such as log Mel-filter bank (FBANK) values. Recent studies found that AMs with convolutional neural networks…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-10 Patrick von Platen , Chao Zhang , Philip Woodland

The past decade has seen substantial work on the use of non-negative matrix factorization and its probabilistic counterparts for audio source separation. Although able to capture audio spectral structure well, these models neglect the…

Machine Learning · Computer Science 2012-07-03 Gautham Mysore , Maneesh Sahani

Voice Activity Detection (VAD) plays a key role in speech processing, often utilizing hand-crafted or neural features. This study examines the effectiveness of Mel-Frequency Cepstral Coefficients (MFCCs) and pre-trained model (PTM)…

Sound · Computer Science 2025-06-03 Kumud Tripathi , Chowdam Venkata Kumar , Pankaj Wasnik

Automated detection of voice disorders with computational methods is a recent research area in the medical domain since it requires a rigorous endoscopy for the accurate diagnosis. Efficient screening methods are required for the diagnosis…

Quantitative Methods · Quantitative Biology 2018-12-06 Vibhuti Gupta

Speech emotion recognition (SER) is a field that has drawn a lot of attention due to its applications in diverse fields. A current trend in methods used for SER is to leverage embeddings from pre-trained models (PTMs) as input features to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Orchid Chetia Phukan , Arun Balaji Buduru , Rajesh Sharma

This paper evaluates the robustness of a DNN-HMM-based speech recognition system in highly-reverberant real environments using the HRRE database. The performance of locally-normalized filter bank (LNFB) and Mel filter bank (MelFB) features…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-28 José Novoa , Juan Pablo Escudero , Jorge Wuth , Victor Poblete , Simon King , Richard Stern , Néstor Becerra Yoma

This work adapts two recent architectures of generative models and evaluates their effectiveness for the conversion of whispered speech to normal speech. We incorporate the normal target speech into the training criterion of…

Speaker verification (SV) utilizing features obtained from models pre-trained via self-supervised learning has recently demonstrated impressive performances. However, these pre-trained models (PTMs) usually have a temporal resolution of 20…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Jisoo Myoung , Sangwook Han , Kihyuk Kim , Jong Won Shin

In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-22 Andrew Huang , Puwei Bao

There are a variety of features of the human voice that can be classified as pitch, timbre, loudness, and vocal tone. It is observed in numerous incidents that human expresses their feelings using different vocal qualities when they are…

Hidden semi-Markov models (HSMMs) are latent variable models which allow latent state persistence and can be viewed as a generalization of the popular hidden Markov models (HMMs). In this paper, we introduce a novel spectral algorithm to…

Machine Learning · Statistics 2016-03-01 Igor Melnyk , Arindam Banerjee