中文
相关论文

相关论文: Dynamic Fusion Multimodal Network for SpeechWellne…

200 篇论文

The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal…

机器学习 · 计算机科学 2023-06-07 Qingyang Zhang , Haitao Wu , Changqing Zhang , Qinghua Hu , Huazhu Fu , Joey Tianyi Zhou , Xi Peng

This study focuses on how different modalities of human communication can be used to distinguish between healthy controls and subjects with schizophrenia who exhibit strong positive symptoms. We developed a multi-modal schizophrenia…

信号处理 · 电气工程与系统科学 2024-04-22 Gowtham Premananth , Yashish M. Siriwardena , Philip Resnik , Carol Espy-Wilson

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025…

Intelligent fault diagnosis has become an indispensable technique for ensuring machinery reliability. However, existing methods suffer significant performance decline in real-world scenarios where models are tested under unseen working…

人工智能 · 计算机科学 2026-01-01 Pengcheng Xia , Yixiang Huang , Chengjin Qin , Chengliang Liu

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine…

Robot-assisted neurological surgery is receiving growing interest due to the improved dexterity, precision, and control of surgical tools, which results in better patient outcomes. However, such systems often limit surgeons' natural sensory…

信号处理 · 电气工程与系统科学 2025-08-13 Zacharias Chen , Alexa Cristelle Cahilig , Sarah Dias , Prithu Kolar , Ravi Prakash , Patrick J. Codd

Leveraging multimodal information from Magnetic Resonance Imaging (MRI) plays a vital role in lesion segmentation, especially for brain tumors. However, in clinical practice, multimodal MRI data are often incomplete, making it challenging…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Yulong Zou , Bo Liu , Cun-Jing Zheng , Yuan-ming Geng , Siyue Li , Qiankun Zuo , Shuihua Wang , Yudong Zhang , Jin Hong

Previous studies have proven that integrating video signals, as a complementary modality, can facilitate improved performance for speech enhancement (SE). However, video clips usually contain large amounts of data and pose a high cost in…

音频与语音处理 · 电气工程与系统科学 2020-08-25 Cheng Yu , Kuo-Hsuan Hung , Syu-Siang Wang , Szu-Wei Fu , Yu Tsao , Jeih-weih Hung

Early detection of cognitive disorders such as Alzheimer's disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce CogniAlign, a multimodal architecture for Alzheimer's…

The dynamic core hypothesis posits that consciousness is correlated with simultaneously integrated and differentiated assemblies of transiently synchronized brain regions. We represented time-dependent functional interactions using dynamic…

神经元与认知 · 定量生物学 2024-06-19 Sofia Morena del Pozo , Helmut Laufs , Vincent Bonhomme , Steven Laureys , Pablo Balenzuela , Enzo Tagliazucchi

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

We propose a novel Dynamic Restrained Uncertainty Weighting Loss to experimentally handle the problem of balancing the contributions of multiple tasks on the ICML ExVo 2022 Challenge. The multitask aims to recognize expressed emotions and…

Survival prediction plays a crucial role in assisting clinicians with the development of cancer treatment protocols. Recent evidence shows that multimodal data can help in the diagnosis of cancer disease and improve survival prediction.…

图像与视频处理 · 电气工程与系统科学 2023-11-14 Ruiquan Ge , Xiangyang Hu , Rungen Huang , Gangyong Jia , Yaqi Wang , Renshu Gu , Changmiao Wang , Elazab Ahmed , Linyan Wang , Juan Ye , Ye Li

Aims Late diagnosis of Oral Squamous Cell Carcinoma (OSCC) contributes significantly to its high global mortality rate, with over 50\% of cases detected at advanced stages and a 5-year survival rate below 50\% according to WHO statistics.…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Ajo Babu George , Sreehari J R Ajo Babu George , Sreehari J R Ajo Babu George , Sreehari J R

The analysis of multivariate time series data is challenging due to the various frequencies of signal changes that can occur over both short and long terms. Furthermore, standard deep learning models are often unsuitable for such datasets,…

机器学习 · 计算机科学 2023-06-21 Iman Deznabi , Madalina Fiterau

Depression commonly co-occurs with neurodegenerative disorders like Multiple Sclerosis (MS), yet the potential of speech-based Artificial Intelligence for detecting depression in such contexts remains unexplored. This study examines the…

Depression remains widely underdiagnosed and undertreated because stigma and subjective symptom ratings hinder reliable screening. To address this challenge, we propose a coarse-to-fine, multi-stage framework that leverages large language…

人工智能 · 计算机科学 2026-04-14 Shiyu Teng , Jiaqing Liu , Hao Sun , Yu Li , Shurong Chai , Ruibo Hou , Tomoko Tateyama , Lanfen Lin , Yen-Wei Chen

Alzheimer's dementia (AD) affects memory, thinking, and language, deteriorating person's life. An early diagnosis is very important as it enables the person to receive medical help and ensure quality of life. Therefore, leveraging…

机器学习 · 计算机科学 2023-04-06 Michail Chatzianastasis , Loukas Ilias , Dimitris Askounis , Michalis Vazirgiannis