中文
相关论文

相关论文: A Comprehensive Rubric for Annotating Pathological…

200 篇论文

The Simon Fraser University Speech Error Database (SFUSED) is a public data collection developed for linguistic and psycholinguistic research. Here we demonstrate how its design and annotations can be used to test and evaluate speech…

计算与语言 · 计算机科学 2025-08-19 John Alderete , Macarious Kin Fung Hui , Aanchan Mohan

Morality plays an important role in culture, identity, and emotion. Recent advances in natural language processing have shown that it is possible to classify moral values expressed in text at scale. Morality classification relies on human…

计算与语言 · 计算机科学 2022-10-17 Negar Mokhberian , Frederic R. Hopp , Bahareh Harandizadeh , Fred Morstatter , Kristina Lerman

Automated audio captioning (AAC) is an important cross-modality translation task, aiming at generating descriptions for audio clips. However, captions generated by previous AAC models have faced ``false-repetition'' errors due to the…

音频与语音处理 · 电气工程与系统科学 2023-06-21 Hanxue Zhang , Zeyu Xie , Xuenan Xu , Mengyue Wu , Kai Yu

Speech pauses, alongside content and structure, offer a valuable and non-invasive biomarker for detecting dementia. This work investigates the use of pause-enriched transcripts in transformer-based language models to differentiate the…

音频与语音处理 · 电气工程与系统科学 2024-08-28 Franziska Braun , Sebastian P. Bayerl , Florian Hönig , Hartmut Lehfeld , Thomas Hillemacher , Tobias Bocklet , Korbinian Riedhammer

This work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine…

音频与语音处理 · 电气工程与系统科学 2022-06-20 Ilja Baumann , Dominik Wagner , Sebastian Bayerl , Tobias Bocklet

Autism Spectrum Disorders (ASD) describe a heterogeneous set of conditions classified as neurodevelopmental disorders. Although the mechanisms underlying ASD are not yet fully understood, more recent literature focused on multiple genetics…

信号处理 · 电气工程与系统科学 2025-01-31 Jessica Vacca , Natascia Brondino , Fabio Dell'Acqua , Anna Vizziello , Pietro Savazzi

In the following report we propose pipelines for Goodness of Pronunciation (GoP) computation solving OOV problem at testing time using Vocab/Lexicon expansion techniques. The pipeline uses different components of ASR system to quantify…

计算与语言 · 计算机科学 2023-03-03 Ankit Grover

As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the…

人工智能 · 计算机科学 2026-03-04 Qiyuan Zhang , Junyi Zhou , Yufei Wang , Fuyuan Lyu , Yidong Ming , Can Xu , Qingfeng Sun , Kai Zheng , Peng Kang , Xue Liu , Chen Ma

The automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Language (L2) learning or…

音频与语音处理 · 电气工程与系统科学 2023-11-14 Mostafa Shahin , Julien Epps , Beena Ahmed

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human…

We propose a model to obtain phonemic and prosodic labels of speech that are coherent with graphemes. Unlike previous methods that simply fine-tune a pre-trained ASR model with the labels, the proposed model conditions the label generation…

声音 · 计算机科学 2025-06-06 Hien Ohnaka , Yuma Shirahata , Byeongseon Park , Ryuichi Yamamoto

Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-speech samples with subjective labels. This study looks at the…

声音 · 计算机科学 2022-09-15 Michael Chinen , Jan Skoglund , Chandan K A Reddy , Alessandro Ragano , Andrew Hines

Background: Subtle changes in spontaneous language production are among the earliest indicators of cognitive decline. Identifying linguistically interpretable markers of dementia can support transparent and clinically grounded screening…

计算与语言 · 计算机科学 2026-02-12 Artsvik Avetisyan , Sachin Kumar

This paper describes an online algorithm for enhancing monaural noisy speech. Firstly, a novel phase-corrected low-delay gammatone filterbank is derived for signal subband decomposition and resynthesis; the subband signals are then analyzed…

声音 · 计算机科学 2015-07-09 Zhangli Chen , Volker Hohmann

Today, data collection has improved in various areas, and the medical domain is no exception. Auscultation, as an important diagnostic technique for physicians, due to the progress and availability of digital stethoscopes, lends itself well…

Voice disorders are pathologies significantly affecting patient quality of life. However, non-invasive automated diagnosis of these pathologies is still under-explored, due to both a shortage of pathological voice data, and diversity of the…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Alkis Koudounas , Gabriele Ciravegna , Marco Fantini , Giovanni Succo , Erika Crosetti , Tania Cerquitelli , Elena Baralis

The detection of perceived prominence in speech has attracted approaches ranging from the design of linguistic knowledge-based acoustic features to the automatic feature learning from suprasegmental attributes such as pitch and intensity…

计算与语言 · 计算机科学 2021-10-28 Mithilesh Vaidya , Kamini Sabu , Preeti Rao

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

音频与语音处理 · 电气工程与系统科学 2025-11-05 Cedric Chan , Jianjing Kuang

Psoriasis (PsO) severity scoring is important for clinical trials but is hindered by inter-rater variability and the burden of in person clinical evaluation. Remote imaging using patient captured mobile photos offers scalability but…

The performances of the automatic speaker verification (ASV) systems degrade due to the reduction in the amount of speech used for enrollment and verification. Combining multiple systems based on different features and classifiers…

计算机视觉与模式识别 · 计算机科学 2019-02-01 Arnab Poddar , Md Sahidullah , Goutam Saha