中文
相关论文

相关论文: Comparative Evaluation of Acoustic Feature Extract…

200 篇论文

Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets,…

音频与语音处理 · 电气工程与系统科学 2024-08-02 Michał Junczyk

In this paper, we present a new open source toolkit for automatic speech recognition (ASR), named CAT (CRF-based ASR Toolkit). A key feature of CAT is discriminative training in the framework of conditional random field (CRF), particularly…

机器学习 · 计算机科学 2019-11-21 Keyu An , Hongyu Xiang , Zhijian Ou

We present our submission to the ICASSP-SPGC-2023 ADReSS-M Challenge Task, which aims to investigate which acoustic features can be generalized and transferred across languages for Alzheimer's Disease (AD) prediction. The challenge consists…

计算与语言 · 计算机科学 2025-01-14 Xuchu Chen , Yu Pu , Jinpeng Li , Wei-Qiang Zhang

Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable caseloads. We test a hierarchical approach to SSD classification on the granular…

计算与语言 · 计算机科学 2026-04-30 Darren Fürst , Sebastian Steindl , Ulrich Schäfer

This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we evaluate six…

We present the Open ASR Leaderboard, a reproducible benchmarking platform with community contributions from academia and industry. It compares 86 open-source and proprietary systems across 12 datasets, with English short- and long-form and…

Background: Alzheimer's disease and related dementias (ADRD) are progressive neurodegenerative conditions where early detection is vital for timely intervention and care. Spontaneous speech contains rich acoustic and linguistic markers that…

计算与语言 · 计算机科学 2025-06-16 Jingyu Li , Lingchao Mao , Hairong Wang , Zhendong Wang , Xi Mao , Xuelei Sherry Ni

Shared decision-making (SDM) is necessary to achieve patient-centred care. Currently no methodology exists to automatically measure SDM at scale. This study aimed to develop an automated approach to measure SDM by using language modelling…

This paper presents Praaline, an open-source software system for managing, annotating, analysing and visualising speech corpora. Researchers working with speech corpora are often faced with multiple tools and formats, and they need to work…

计算与语言 · 计算机科学 2018-02-09 George Christodoulides

Voice biomarkers--human-generated acoustic signals such as speech, coughing, and breathing--are promising tools for scalable, non-invasive detection and monitoring of mental health and neurodegenerative diseases. Yet, their clinical…

声音 · 计算机科学 2025-08-21 Ishaan Mahapatra , Nihar R. Mahapatra

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and…

This study investigates factors influencing Automatic Speech Recognition (ASR) systems' fairness and performance across genders, beyond the conventional examination of demographics. Using the LibriSpeech dataset and the Whisper small model,…

计算与语言 · 计算机科学 2025-02-26 Hend ElGhazaly , Bahman Mirheidari , Nafise Sadat Moosavi , Heidi Christensen

Recently, deep neural network (DNN)-based speech enhancement (SE) systems have been used with great success. During training, such systems require clean speech data - ideally, in large quantity with a variety of acoustic conditions, many…

音频与语音处理 · 电气工程与系统科学 2021-05-27 Koichi Saito , Stefan Uhlich , Giorgio Fabbro , Yuki Mitsufuji

Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable…

声音 · 计算机科学 2024-12-06 Yerin Choi , Jeehyun Lee , Myoung-Wan Koo

Computer assisted technologies based on algorithmic software segmentation are an increasing topic of interest in complex surgical cases. However - due to functional instability, time consuming software processes, personnel resources or…

Advances in artificial intelligence (AI) and deep learning have improved diagnostic capabilities in healthcare, yet limited interpretability continues to hinder clinical adoption. Schizophrenia, a complex disorder with diverse symptoms…

音频与语音处理 · 电气工程与系统科学 2025-11-06 Gowtham Premananth , Carol Espy-Wilson

Social determinants of health (SDoH) have an important impact on patient outcomes but are incompletely collected from the electronic health records (EHR). This study researched the ability of large language models to extract SDoH from free…

Speech enhancement techniques improve the quality or the intelligibility of an audio signal by removing unwanted noise. It is used as preprocessing in numerous applications such as speech recognition, hearing aids, broadcasting and…

音频与语音处理 · 电气工程与系统科学 2023-06-05 Angélica S. Z. Suárez , Clément Laroche , Line H. Clemmensen , Sneha Das

Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful application of best…

声音 · 计算机科学 2025-08-07 Eduardo Pacheco , Atila Orhon , Berkin Durmus , Blaise Munyampirwa , Andrey Leonov

Audio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning-based codecs are often tested on datasets with similar…

音频与语音处理 · 电气工程与系统科学 2025-08-04 Jozef Coldenhoff , Niclas Granqvist , Milos Cernak