中文
相关论文

相关论文: Multimodal Magic Elevating Depression Detection wi…

200 篇论文

Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose a multi-agent reinforcement learning framework that…

人工智能 · 计算机科学 2025-10-14 Jiabao Shi , Minfeng Qi , Lefeng Zhang , Di Wang , Yingjie Zhao , Ziying Li , Yalong Xing , Ningran Li

In recent years, researchers combine both audio and video signals to deal with challenges where actions are not well represented or captured by visual cues. However, how to effectively leverage the two modalities is still under development.…

计算机视觉与模式识别 · 计算机科学 2024-01-09 Wentao Zhu

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Xiangyu Zhang , Hexin Liu , Kaishuai Xu , Qiquan Zhang , Daijiao Liu , Beena Ahmed , Julien Epps

Multi-modal fusion methods often suffer from two types of representation collapse: feature collapse where individual dimensions lose their discriminative power (as measured by eigenspectra), and modality collapse where one dominant modality…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Seulgi Kim , Kiran Kokilepersaud , Mohit Prabhushankar , Ghassan AlRegib

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

音频与语音处理 · 电气工程与系统科学 2026-01-23 Ina Salaj , Arijit Biswas

Though multimodal emotion recognition has achieved significant progress over recent years, the potential of rich synergic relationships across the modalities is not fully exploited. In this paper, we introduce Recursive Joint Cross-Modal…

计算机视觉与模式识别 · 计算机科学 2024-04-16 R. Gnana Praveen , Jahangir Alam

In recent years, multi-modal fusion has attracted a lot of research interest, both in academia, and in industry. Multimodal fusion entails the combination of information from a set of different types of sensors. Exploiting complementary…

机器学习 · 计算机科学 2020-08-27 Siddharth Roheda , Hamid Krim , Benjamin S. Riggan

Discovering materials with desirable properties in an efficient way remains a significant problem in materials science. Many studies have tackled this problem by using different sets of information available about the materials. Among them,…

材料科学 · 物理学 2025-03-04 Onur Boyar , Indra Priyadarsini , Seiji Takeda , Lisa Hamada

The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Luyao Lei , Shuo Xu , Yifan Bai , Xing Wei

Deep learning has proven to be a highly effective tool for a wide range of applications, significantly when leveraging the power of multi-loss functions to optimize performance on multiple criteria simultaneously. However, optimal selection…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Amin Golnari , Mostafa Diba

Recently, learning-based approaches show promising results in navigation tasks. However, the poor generalization capability and the simulation-reality gap prevent a wide range of applications. We consider the problem of improving the…

机器人学 · 计算机科学 2023-09-26 Wenzhe Cai , Guangran Cheng , Lingyue Kong , Lu Dong , Changyin Sun

Multimodal emotion recognition aims to recognize emotions for each utterance of multiple modalities, which has received increasing attention for its application in human-machine interaction. Current graph-based methods fail to…

计算与语言 · 计算机科学 2023-11-21 Dongyuan Li , Yusong Wang , Kotaro Funakoshi , Manabu Okumura

Body-conduction microphone signals (BMS) bypass airborne sound, providing strong noise resistance. However, a complementary modality is required to compensate for the inherent loss of high-frequency information. In this study, we propose a…

声音 · 计算机科学 2025-08-29 Yunsik Kim , Yoonyoung Chung

Bipolar disorder is a mental health disorder that causes mood swings that range from depression to mania. Diagnosis of bipolar disorder is usually done based on patient interviews, and reports obtained from the caregivers of the patients.…

计算与语言 · 计算机科学 2021-12-20 Pınar Baki

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Haotian Wang , Yuxuan Xi , Hang Chen , Jun Du , Yan Song , Qing Wang , Hengshun Zhou , Chenxi Wang , Jiefeng Ma , Pengfei Hu , Ya Jiang , Shi Cheng , Jie Zhang , Yuzhe Weng

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

音频与语音处理 · 电气工程与系统科学 2025-01-16 Amit Eliav , Sharon Gannot

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

In the past, Acoustic Scene Classification systems have been based on hand crafting audio features that are input to a classifier. Nowadays, the common trend is to adopt data driven techniques, e.g., deep learning, where audio…

声音 · 计算机科学 2018-06-29 Eduardo Fonseca , Rong Gong , Xavier Serra

Multimodal Emotion Recognition (MER) aims to perceive human emotions through three modes: language, vision, and audio. Previous methods primarily focused on modal fusion without adequately addressing significant distributional differences…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Jichao Zhu , Jun Yu

As a sub-branch of affective computing, impression recognition, e.g., perception of speaker characteristics such as warmth or competence, is potentially a critical part of both human-human conversations and spoken dialogue systems. Most…

多媒体 · 计算机科学 2023-02-17 Yuanchao Li , Peter Bell , Catherine Lai
‹ 上一页 1 8 9 10 下一页 ›