中文
相关论文

相关论文: Multimodal Audio-based Disease Prediction with Tra…

200 篇论文

Alzheimer's Disease (AD) is a complex neurodegenerative disorder marked by memory loss, executive dysfunction, and personality changes. Early diagnosis is challenging due to subtle symptoms and varied presentations, often leading to…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Yifei Chen , Shenghao Zhu , Zhaojie Fang , Chang Liu , Binfeng Zou , Yuhe Wang , Shuo Chang , Fan Jia , Feiwei Qin , Jin Fan , Yong Peng , Changmiao Wang

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Wangyuan Zhu , Jun Yu

Depression is a serious mental health illness that significantly affects an individual's well-being and quality of life, making early detection crucial for adequate care and treatment. Detecting depression is often difficult, as it is based…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Md Rezwanul Haque , Md. Milon Islam , S M Taslim Uddin Raju , Hamdi Altaheri , Lobna Nassar , Fakhri Karray

Alzheimer's disease (AD) is a progressive neurodegenerative disease with high inter-patient variance in rate of cognitive decline. AD progression prediction aims to forecast patient cognitive decline and benefits from incorporating multiple…

机器学习 · 计算机科学 2025-09-17 Benjamin Burns , Yuan Xue , Douglas W. Scharre , Xia Ning

Detecting collaborative and problem-solving behaviours from digital traces to interpret students' collaborative problem solving (CPS) competency is a long-term goal in the Artificial Intelligence in Education (AIEd) field. Although…

计算与语言 · 计算机科学 2025-04-22 K. Wong , B. Wu , S. Bulathwela , M. Cukurova

Human-machine interaction has been around for several decades now, with new applications emerging every day. One of the major goals that remain to be achieved is designing an interaction similar to how a human interacts with another human.…

人机交互 · 计算机科学 2022-12-27 Tauheed Khan Mohd , Nicole Nguyen , Ahmad Y Javaid

With the increasing availability of diverse data types, particularly images and time series data from medical experiments, there is a growing demand for techniques designed to combine various modalities of data effectively. Our motivation…

图像与视频处理 · 电气工程与系统科学 2024-05-27 Ali Rasekh , Reza Heidari , Amir Hosein Haji Mohammad Rezaie , Parsa Sharifi Sedeh , Zahra Ahmadi , Prasenjit Mitra , Wolfgang Nejdl

Multimodal research is an emerging field of artificial intelligence, and one of the main research problems in this field is multimodal fusion. The fusion of multimodal data is the process of integrating multiple unimodal representations…

Accurate diagnosis of Alzheimer's disease (AD) is essential for enabling timely intervention and slowing disease progression. Multimodal diagnostic approaches offer considerable promise by integrating complementary information across…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Yujie Nie , Jianzhang Ni , Yonglong Ye , Yuan-Ting Zhang , Yun Kwok Wing , Xiangqing Xu , Xin Ma , Lizhou Fan

Recent technological advances in healthcare have led to unprecedented growth in patient data quantity and diversity. While artificial intelligence (AI) models have shown promising results in analyzing individual data modalities, there is…

Gaining insights into the structural and functional mechanisms of the brain has been a longstanding focus in neuroscience research, particularly in the context of understanding and treating neuropsychiatric disorders such as Schizophrenia…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Badhan Mazumder , Lei Wu , Vince D. Calhoun , Dong Hye Ye

Speech is a noninvasive digital phenotype that can offer valuable insights into mental health conditions, but it is often treated as a single modality. In contrast, we propose the treatment of patient speech data as a trimodal multimedia…

The current clinical diagnosis framework of Alzheimer's disease (AD) involves multiple modalities acquired from multiple diagnosis stages, each with distinct usage and cost. Previous AD diagnosis research has predominantly focused on how to…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Yuxiao Liu , Mianxin Liu , Yuanwang Zhang , Kaicong Sun , Dinggang Shen

We propose a multimodal latent diffusion model that jointly synthesizes volumetric magnetic resonance imaging (MRI) and tabular clinical data within a shared latent space via cross-attention. This approach enables coherent joint…

图像与视频处理 · 电气工程与系统科学 2026-05-11 Daniel Mensing , Jan Kapar , Jochen G. Hirsch , Matthias Günther , Horst Hahn , Marvin N. Wright

AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We…

声音 · 计算机科学 2025-08-05 Fan Wu , Kaicheng Zhao , Elgar Fleisch , Filipe Barata

Managing fluid balance in dialysis patients is crucial, as improper management can lead to severe complications. In this paper, we propose a multimodal approach that integrates visual features from lung ultrasound images with clinical data…

图像与视频处理 · 电气工程与系统科学 2024-10-04 Tianqi Yang , Nantheera Anantrasirichai , Oktay Karakuş , Marco Allinovi , Alin Achim

Deep learning and multi-modal fusion have demonstrated transformative potential in medical diagnosis by integrating diverse data sources. However, accurate prognosis for ischemic stroke remains challenging due to limitations in existing…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Liren Chen , Lidong Sun , Mingyan Huang , Junzhe Tang , Yinghui Zhu , Guanjie Wang , Yiqing Xia , Ting Xiao

In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals…

声音 · 计算机科学 2026-01-14 Tiantian Feng , Anfeng Xu , Jinkook Lee , Shrikanth Narayanan

Clinical diagnostic decision making and population-based studies often rely on multi-modal data which is noisy and incomplete. Recently, several works proposed geometric deep learning approaches to solve disease classification, by modeling…

机器学习 · 计算机科学 2019-05-09 Gerome Vivar , Hendrik Burwinkel , Anees Kazi , Andreas Zwergal , Nassir Navab , Seyed-Ahmad Ahmadi

In this study, we present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesize that explicitly disentangling shared and modality-specific representations within multimodal…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Mohd Mujtaba Akhtar , Girish , Muskaan Singh