English
Related papers

Related papers: Optimising MFCC parameters for the automatic detec…

200 papers

A novel text-independent speaker identification (SI) method is proposed. This method uses the Mel-frequency Cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature sets to capture speaker's…

Sound · Computer Science 2020-02-04 Zhanyu Ma , Hong Yu

Maximum Voiced Frequency (MVF) is used in various speech models as the spectral boundary separating periodic and aperiodic components during the production of voiced sounds. Recent studies have shown that its proper estimation and modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-02 Thomas Drugman , Yannis Stylianou

Recent advancements in deep learning techniques have sparked performance boosts in various real-world applications including disease diagnosis based on multi-modal medical data. Cough sound data-based respiratory disease (e.g., COVID-19 and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Qian Wang , Zhaoyang Bu , Jiaxuan Mao , Wenyu Zhu , Jingya Zhao , Wei Du , Guochao Shi , Min Zhou , Si Chen , Jieming Qu

In this paper, a data augmentation method is proposed for depression detection from speech signals. Samples for data augmentation were created by changing the frame-width and the frame-shift parameters during the feature extraction process.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-15 Vijay Ravi , Jinhan Wang , Jonathan Flint , Abeer Alwan

Respiratory rate (RR) is a clinical metric used to assess overall health and physical fitness. An individual's RR can change from their baseline due to chronic illness symptoms (e.g., asthma, congestive heart failure), acute illness (e.g.,…

Sound · Computer Science 2021-07-30 Agni Kumar , Vikramjit Mitra , Carolyn Oliver , Adeeti Ullal , Matt Biddulph , Irida Mance

Availability of diagnostic codes in Electronic Health Records (EHRs) is crucial for patient care as well as reimbursement purposes. However, entering them in the EHR is tedious, and some clinical codes may be overlooked. Given an…

Machine Learning · Computer Science 2023-05-10 Tsvetan R. Yordanov , Ameen Abu-Hanna , Anita CJ Ravelli , Iacopo Vagliano

Accurate, fast, and reliable multiclass classification of electroencephalography (EEG) signals is a challenging task towards the development of motor imagery brain-computer interface (MI-BCI) systems. We propose enhancements to different…

Signal Processing · Electrical Eng. & Systems 2018-12-14 Michael Hersche , Tino Rellstab , Pasquale Davide Schiavone , Lukas Cavigelli , Luca Benini , Abbas Rahimi

The early and accurate diagnosis of sepsis is critical for enhancing patient outcomes. This study aims to use heart rate variability (HRV) features to develop an effective predictive model for sepsis detection. Critical HRV features are…

Machine Learning · Computer Science 2024-08-07 Sai Balaji , Christopher Sun , Anaiy Somalwar

The COVID-19 pandemic has accelerated research on design of alternative, quick and effective COVID-19 diagnosis approaches. In this paper, we describe the Coswara tool, a website application designed to enable COVID-19 detection by…

The purpose of speech emotion recognition system is to classify speakers utterances into different emotional states such as disgust, boredom, sadness, neutral and happiness. Speech features that are commonly used in speech emotion…

Computation and Language · Computer Science 2014-06-25 Imen Trabelsi , Dorra Ben Ayed , Noureddine Ellouze

The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech…

Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improve performance of Automatic Speech Recognition (ASR) systems,…

In this work, we propose an ensemble of classifiers to distinguish between various degrees of abnormalities of the heart using Phonocardiogram (PCG) signals acquired using digital stethoscopes in a clinical setting, for the INTERSPEECH 2018…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Ahmed Imtiaz Humayun , Md. Tauhiduzzaman Khan , Shabnam Ghaffarzadegan , Zhe Feng , Taufiq Hasan

The issue in respiratory sound classification has attained good attention from the clinical scientists and medical researcher's group in the last year to diagnosing COVID-19 disease. To date, various models of Artificial Intelligence (AI)…

Sound · Computer Science 2021-12-15 Kranthi Kumar Lella , Alphonse Pja

There are different algorithms for vocal fold pathology diagnosis. These algorithms usually have three stages which are Feature Extraction, Feature Reduction and Classification. While the third stage implies a choice of a variety of machine…

Machine Learning · Computer Science 2013-02-08 Vahid Majidnezhad , Igor Kheidorov

Voice activity detection (VAD), used as the front end of speech enhancement, speech and speaker recognition algorithms, determines the overall accuracy and efficiency of the algorithms. Therefore, a VAD with low complexity and high accuracy…

Sound · Computer Science 2019-02-06 Jayanta Dey , Md Sanzid Bin Hossain , Mohammad Ariful Haque

In Acoustic Scene Classification (ASC) two major approaches have been followed . While one utilizes engineered features such as mel-frequency-cepstral-coefficients (MFCCs), the other uses learned features that are the outcome of an…

Sound · Computer Science 2017-11-15 Hamid Eghbal-zadeh , Bernhard Lehner , Matthias Dorfer , Gerhard Widmer

As respiratory illnesses become more common, it is crucial to quickly and accurately detect them to improve patient care. There is a need for improved diagnostic methods for immediate medical assessments for optimal patient outcomes. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-30 Paridhi Mundra , Manik Sharma , Yashwardhan Chaudhuri , Orchid Chetia Phukan , Arun Balaji Buduru

This paper evaluates a wide range of audio-based deep learning frameworks applied to the breathing, cough, and speech sounds for detecting COVID-19. In general, the audio recording inputs are transformed into low-level spectrogram features,…

Sound · Computer Science 2022-03-03 Dat Ngo , Lam Pham , Truong Hoang , Sefki Kolozali , Delaram Jarchi

This work is devoted to capturing Emirati-accented speech database (Arabic United Arab Emirates database) in each of neutral and shouted talking environments in order to study and enhance text-independent Emirati-accented speaker…

Sound · Computer Science 2018-04-04 Ismail Shahin , Ali Bou Nassif , Mohammed Bahutair