中文
相关论文

相关论文: Temporal-Aware Iterative Speech Model for Dementia…

200 篇论文

\textbf{Background:} Machine learning models trained on electronic health records (EHRs) often degrade across healthcare systems due to distributional shift. A fundamental but underexplored factor is diagnostic signal decay: variability in…

机器学习 · 计算机科学 2025-09-11 Jingya Cheng , Jiazi Tian , Federica Spoto , Alaleh Azhir , Daniel Mork , Hossein Estiri

Alzheimer's Disease (AD) is a severe brain disorder, destroying memories and brain functions. AD causes chronically, progressively, and irreversibly cognitive declination and brain damages. The reliable and effective evaluation of early…

图像与视频处理 · 电气工程与系统科学 2021-01-07 Kuo Yang , Emad A. Mohammed

Multi-modal biological, imaging, and neuropsychological markers have demonstrated promising performance for distinguishing Alzheimer's disease (AD) patients from cognitively normal elders. However, it remains difficult to early predict when…

计算机视觉与模式识别 · 计算机科学 2019-01-08 Hongming Li , Yong Fan

Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care to delay further progression. This paper presents the development of a state-of-the-art Conformer based speech recognition system built on the…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Tianzi Wang , Jiajun Deng , Mengzhe Geng , Zi Ye , Shoukang Hu , Yi Wang , Mingyu Cui , Zengrui Jin , Xunying Liu , Helen Meng

Early detection of dementia is critical for timely medical intervention and improved patient outcomes. Neuropsychological tests are widely used for cognitive assessment but have traditionally relied on manual scoring. Automatic dementia…

机器学习 · 计算机科学 2025-07-15 Liming Wang , Saurabhchand Bhati , Cody Karjadi , Rhoda Au , James Glass

Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care and delay progression. Speech based automatic AD screening systems provide a non-intrusive and more scalable alternative to other clinical screening…

机器学习 · 计算机科学 2022-08-09 Yi Wang , Tianzi Wang , Zi Ye , Lingwei Meng , Shoukang Hu , Xixin Wu , Xunying Liu , Helen Meng

This paper introduces an improved duration informed attention neural network (DurIAN-E) for expressive and high-fidelity text-to-speech (TTS) synthesis. Inherited from the original DurIAN model, an auto-regressive model structure in which…

音频与语音处理 · 电气工程与系统科学 2023-09-25 Yu Gu , Yianrao Bian , Guangzhi Lei , Chao Weng , Dan Su

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Nowadays, a huge number of images are available. However, retrieving a required image for an ordinary user is a challenging task in computer vision systems. During the past two decades, many types of research have been introduced to improve…

多媒体 · 计算机科学 2020-01-30 Amir Vatani , Milad Taleby Ahvanooey , Mostafa Rahimi

Speech emotion recognition (SER) plays a vital role in improving the interactions between humans and machines by inferring human emotion and affective states from speech signals. Whereas recent works primarily focus on mining spatiotemporal…

声音 · 计算机科学 2023-10-03 Jiaxin Ye , Xin-cheng Wen , Yujie Wei , Yong Xu , Kunhong Liu , Hongming Shan

The ADReSS Challenge at INTERSPEECH 2020 defines a shared task through which different approaches to the automated recognition of Alzheimer's dementia based on spontaneous speech can be compared. ADReSS provides researchers with a benchmark…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Saturnino Luz , Fasih Haider , Sofia de la Fuente , Davida Fromm , Brian MacWhinney

Automatic recognition of disordered speech remains a highly challenging task to date. Sources of variability commonly found in normal speech including accent, age or gender, when further compounded with the underlying causes of speech…

声音 · 计算机科学 2022-01-20 Mengzhe Geng , Shansong Liu , Jianwei Yu , Xurong Xie , Shoukang Hu , Zi Ye , Zengrui Jin , Xunying Liu , Helen Meng

We present two multimodal fusion-based deep learning models that consume ASR transcribed speech and acoustic data simultaneously to classify whether a speaker in a structured diagnostic task has Alzheimer's Disease and to what degree,…

计算与语言 · 计算机科学 2021-07-01 Morteza Rohanian , Julian Hough , Matthew Purver

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this system is the Duration Informed Attention Network (DurIAN), an…

计算与语言 · 计算机科学 2019-09-09 Chengzhu Yu , Heng Lu , Na Hu , Meng Yu , Chao Weng , Kun Xu , Peng Liu , Deyi Tuo , Shiyin Kang , Guangzhi Lei , Dan Su , Dong Yu

Speech intelligibility assessment plays an important role in the therapy of patients suffering from pathological speech disorders. Automatic and objective measures are desirable to assist therapists in their traditionally subjective and…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Tobias Weise , Philipp Klumpp , Kubilay Can Demir , Andreas Maier , Elmar Noeth , Bjoern Heismann , Maria Schuster , Seung Hee Yang

The ageing population trend is correlated with an increased prevalence of acquired cognitive impairments such as dementia. Although there is no cure for dementia, a timely diagnosis helps in obtaining necessary support and appropriate…

计算机视觉与模式识别 · 计算机科学 2021-01-06 Xing Liang , Anastassia Angelopoulou , Epaminondas Kapetanios , Bencie Woll , Reda Al-batat , Tyron Woolfe

Cognitive decline is a sign of Alzheimer's disease (AD), and there is evidence that tracking a person's eye movement, using eye tracking devices, can be used for the automatic identification of early signs of cognitive decline. However,…

Speech activity detection (SAD) plays an important role in current speech processing systems, including automatic speech recognition (ASR). SAD is particularly difficult in environments with acoustic noise. A practical solution is to…

计算与语言 · 计算机科学 2023-05-15 Fei Tao , Carlos Busso

Text-to-image models generate highly realistic images based on natural language descriptions and millions of users use them to create and share images online. While it is expected that such models can align input text and generated image in…

机器学习 · 计算机科学 2025-08-14 Mansi , Anastasios Lepipas , Dominika Woszczyk , Yiying Guan , Soteris Demetriou