中文
相关论文

相关论文: Optimising MFCC parameters for the automatic detec…

200 篇论文

A novel text-independent speaker identification (SI) method is proposed. This method uses the Mel-frequency Cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature sets to capture speaker's…

声音 · 计算机科学 2020-02-04 Zhanyu Ma , Hong Yu

Maximum Voiced Frequency (MVF) is used in various speech models as the spectral boundary separating periodic and aperiodic components during the production of voiced sounds. Recent studies have shown that its proper estimation and modeling…

音频与语音处理 · 电气工程与系统科学 2020-06-02 Thomas Drugman , Yannis Stylianou

Recent advancements in deep learning techniques have sparked performance boosts in various real-world applications including disease diagnosis based on multi-modal medical data. Cough sound data-based respiratory disease (e.g., COVID-19 and…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Qian Wang , Zhaoyang Bu , Jiaxuan Mao , Wenyu Zhu , Jingya Zhao , Wei Du , Guochao Shi , Min Zhou , Si Chen , Jieming Qu

In this paper, a data augmentation method is proposed for depression detection from speech signals. Samples for data augmentation were created by changing the frame-width and the frame-shift parameters during the feature extraction process.…

音频与语音处理 · 电气工程与系统科学 2022-02-15 Vijay Ravi , Jinhan Wang , Jonathan Flint , Abeer Alwan

Respiratory rate (RR) is a clinical metric used to assess overall health and physical fitness. An individual's RR can change from their baseline due to chronic illness symptoms (e.g., asthma, congestive heart failure), acute illness (e.g.,…

声音 · 计算机科学 2021-07-30 Agni Kumar , Vikramjit Mitra , Carolyn Oliver , Adeeti Ullal , Matt Biddulph , Irida Mance

Availability of diagnostic codes in Electronic Health Records (EHRs) is crucial for patient care as well as reimbursement purposes. However, entering them in the EHR is tedious, and some clinical codes may be overlooked. Given an…

机器学习 · 计算机科学 2023-05-10 Tsvetan R. Yordanov , Ameen Abu-Hanna , Anita CJ Ravelli , Iacopo Vagliano

Accurate, fast, and reliable multiclass classification of electroencephalography (EEG) signals is a challenging task towards the development of motor imagery brain-computer interface (MI-BCI) systems. We propose enhancements to different…

信号处理 · 电气工程与系统科学 2018-12-14 Michael Hersche , Tino Rellstab , Pasquale Davide Schiavone , Lukas Cavigelli , Luca Benini , Abbas Rahimi

The early and accurate diagnosis of sepsis is critical for enhancing patient outcomes. This study aims to use heart rate variability (HRV) features to develop an effective predictive model for sepsis detection. Critical HRV features are…

机器学习 · 计算机科学 2024-08-07 Sai Balaji , Christopher Sun , Anaiy Somalwar

The COVID-19 pandemic has accelerated research on design of alternative, quick and effective COVID-19 diagnosis approaches. In this paper, we describe the Coswara tool, a website application designed to enable COVID-19 detection by…

The purpose of speech emotion recognition system is to classify speakers utterances into different emotional states such as disgust, boredom, sadness, neutral and happiness. Speech features that are commonly used in speech emotion…

计算与语言 · 计算机科学 2014-06-25 Imen Trabelsi , Dorra Ben Ayed , Noureddine Ellouze

The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech…

Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improve performance of Automatic Speech Recognition (ASR) systems,…

In this work, we propose an ensemble of classifiers to distinguish between various degrees of abnormalities of the heart using Phonocardiogram (PCG) signals acquired using digital stethoscopes in a clinical setting, for the INTERSPEECH 2018…

计算机视觉与模式识别 · 计算机科学 2019-04-30 Ahmed Imtiaz Humayun , Md. Tauhiduzzaman Khan , Shabnam Ghaffarzadegan , Zhe Feng , Taufiq Hasan

The issue in respiratory sound classification has attained good attention from the clinical scientists and medical researcher's group in the last year to diagnosing COVID-19 disease. To date, various models of Artificial Intelligence (AI)…

声音 · 计算机科学 2021-12-15 Kranthi Kumar Lella , Alphonse Pja

There are different algorithms for vocal fold pathology diagnosis. These algorithms usually have three stages which are Feature Extraction, Feature Reduction and Classification. While the third stage implies a choice of a variety of machine…

机器学习 · 计算机科学 2013-02-08 Vahid Majidnezhad , Igor Kheidorov

Voice activity detection (VAD), used as the front end of speech enhancement, speech and speaker recognition algorithms, determines the overall accuracy and efficiency of the algorithms. Therefore, a VAD with low complexity and high accuracy…

声音 · 计算机科学 2019-02-06 Jayanta Dey , Md Sanzid Bin Hossain , Mohammad Ariful Haque

In Acoustic Scene Classification (ASC) two major approaches have been followed . While one utilizes engineered features such as mel-frequency-cepstral-coefficients (MFCCs), the other uses learned features that are the outcome of an…

声音 · 计算机科学 2017-11-15 Hamid Eghbal-zadeh , Bernhard Lehner , Matthias Dorfer , Gerhard Widmer

As respiratory illnesses become more common, it is crucial to quickly and accurately detect them to improve patient care. There is a need for improved diagnostic methods for immediate medical assessments for optimal patient outcomes. This…

音频与语音处理 · 电气工程与系统科学 2024-07-30 Paridhi Mundra , Manik Sharma , Yashwardhan Chaudhuri , Orchid Chetia Phukan , Arun Balaji Buduru

This paper evaluates a wide range of audio-based deep learning frameworks applied to the breathing, cough, and speech sounds for detecting COVID-19. In general, the audio recording inputs are transformed into low-level spectrogram features,…

声音 · 计算机科学 2022-03-03 Dat Ngo , Lam Pham , Truong Hoang , Sefki Kolozali , Delaram Jarchi

This work is devoted to capturing Emirati-accented speech database (Arabic United Arab Emirates database) in each of neutral and shouted talking environments in order to study and enhance text-independent Emirati-accented speaker…

声音 · 计算机科学 2018-04-04 Ismail Shahin , Ali Bou Nassif , Mohammed Bahutair