English
Related papers

Related papers: Causal-Anticausal Decomposition of Speech using Co…

200 papers

In this paper, we propose a novel family of windowing technique to compute Mel Frequency Cepstral Coefficient (MFCC) for automatic speaker recognition from speech. The proposed method is based on fundamental property of discrete time…

Computer Vision and Pattern Recognition · Computer Science 2015-06-05 Md. Sahidullah , Goutam Saha

This paper presents novel Weighted Finite-State Transducer (WFST) topologies to implement Connectionist Temporal Classification (CTC)-like algorithms for automatic speech recognition. Three new CTC variants are proposed: (1) the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-27 Aleksandr Laptev , Somshubra Majumdar , Boris Ginsburg

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide…

Short-time Fourier transform (STFT) is used as the front end of many popular successful monaural speech separation methods, such as deep clustering (DPCL), permutation invariant training (PIT) and their various variants. Since the frequency…

Sound · Computer Science 2019-02-05 Ziqiang Shi , Huibin Lin , Liu Liu , Rujie Liu , Jiqing Han

Mask-based blind speech separation (BSS) estimates source-wise time-frequency (TF) masks by clustering multichannel observations using spatial information. The directional statistical approach clusters normalized multichannel observations…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Nobutaka Ito

Systolic murmurs are extra heart sounds that occur during the contraction phase of the cardiac cycle, often indicating heart abnormalities caused by turbulent blood flow. Their intensity, pitch, and quality vary, requiring precise…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mahmoud Fakhry , Abeer FathAllah Brery

Pitch and Formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are essentially resonance…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-09 Seyedamiryousef Hosseini Goki , Mahdieh Ghazvini , Sajad Hamzenejadi

With the increasing demand for audio communication and online conference, ensuring the robustness of Acoustic Echo Cancellation (AEC) under the complicated acoustic scenario including noise, reverberation and nonlinear distortion has become…

Sound · Computer Science 2022-02-16 Shimin Zhang , Yuxiang Kong , Shubo Lv , Yanxin Hu , Lei Xie

Causal Disentangled Representation Learning(CDRL) aims to learn and disentangle low dimensional representations and their underlying causal structure from observations. However, existing disentanglement methods rely on a standard mean-field…

Machine Learning · Computer Science 2026-01-30 Yutao Jin , Yuang Tao , Junyong Zhai

In this paper, we present a new method for recognizing tones in continuous speech for tonal languages. The method works by converting the speech signal to a cepstrogram, extracting a sequence of cepstral features using a convolutional…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-09 Loren Lugosch , Vikrant Singh Tomar

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio…

Sound · Computer Science 2024-01-19 Yimin Deng , Huaizhen Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

The driving force behind the recent success of LSTMs has been their ability to learn complex and non-linear relationships. Consequently, our inability to describe these relationships has led to LSTMs being characterized as black boxes. To…

Computation and Language · Computer Science 2018-05-01 W. James Murdoch , Peter J. Liu , Bin Yu

We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with…

Sound · Computer Science 2021-03-17 Hang Lv , Zhehuai Chen , Hainan Xu , Daniel Povey , Lei Xie , Sanjeev Khudanpur

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

Sound · Computer Science 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

The intersection of technology and mental health has spurred innovative approaches to assessing emotional well-being, particularly through computational techniques applied to audio data analysis. This study explores the application of…

Sound · Computer Science 2024-12-17 Idoko Agbo , Dr Hoda El-Sayed , M. D Kamruzzan Sarker

Heterogeneity in medical data, e.g., from data collected at different sites and with different protocols in a clinical study, is a fundamental hurdle for accurate prediction using machine learning models, as such models often fail to…

Machine Learning · Computer Science 2021-07-13 Rongguang Wang , Pratik Chaudhari , Christos Davatzikos

In this paper, we propose a classification based glottal closure instants (GCI) detection from pathological acoustic speech signal, which finds many applications in vocal disorder analysis. Till date, GCI for pathological disorder is…

Sound · Computer Science 2018-11-28 Gurunath Reddy M , Tanumay Mandal , Krothapalli Sreenivasa Rao

Recent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences. Such systems are used with a dictionary and…

Computation and Language · Computer Science 2017-03-23 Kartik Audhkhasi , Bhuvana Ramabhadran , George Saon , Michael Picheny , David Nahamoo

We propose a novel phrase break prediction method that combines implicit features extracted from a pre-trained large language model, a.k.a BERT, and explicit features extracted from BiLSTM with linguistic features. In conventional BiLSTM…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-27 Kosuke Futamata , Byeongseon Park , Ryuichi Yamamoto , Kentaro Tachibana

Chordal and factor-width decomposition methods for semidefinite programming and polynomial optimization have recently enabled the analysis and control of large-scale linear systems and medium-scale nonlinear systems. Chordal decomposition…

Optimization and Control · Mathematics 2021-11-23 Yang Zheng , Giovanni Fantuzzi , Antonis Papachristodoulou