中文
相关论文

相关论文: EMISSOR: A platform for capturing multimodal inter…

200 篇论文

Temporal Knowledge Graph (TKG) reasoning often involves completing missing factual elements along the timeline. Although existing methods can learn good embeddings for each factual element in quadruples by integrating temporal information,…

人工智能 · 计算机科学 2024-05-02 Zhiyu Fang , Shuai-Long Lei , Xiaobin Zhu , Chun Yang , Shi-Xue Zhang , Xu-Cheng Yin , Jingyan Qin

Multimodal physiological signals, such as EEG, ECG, EOG, and EMG, are crucial for healthcare and brain-computer interfaces. While existing methods rely on specialized architectures and dataset-specific fusion strategies, they struggle to…

信号处理 · 电气工程与系统科学 2026-03-18 Wei-Bang Jiang , Xi Fu , Yi Ding , Cuntai Guan

An end-to-end platform assembling multiple tiers is built for precisely cognizing brain activities. Being fed massive electroencephalogram (EEG) data, the time-frequency spectrograms are conventionally projected into the episode-wise…

信号处理 · 电气工程与系统科学 2022-04-22 Zheng Chen , Lingwei Zhu , Ziwei Yang , Renyuan Zhang

Modern urban spaces are equipped with an increasingly diverse set of sensors, all producing an abundance of multimodal data. Such multimodal data can be used to identify and reason about important incidents occurring in urban landscapes,…

人工智能 · 计算机科学 2026-02-18 Brian Wang , Mani Srivastava

Multimodal Sentiment Analysis is an active area of research that leverages multimodal signals for affective understanding of user-generated videos. The predominant approach, addressing this task, has been to develop sophisticated fusion…

计算与语言 · 计算机科学 2020-10-20 Devamanyu Hazarika , Roger Zimmermann , Soujanya Poria

EEG-based multimodal emotion recognition(EMER) has gained significant attention and witnessed notable advancements, the inherent complexity of human neural systems has motivated substantial efforts toward multimodal approaches. However,…

信号处理 · 电气工程与系统科学 2025-10-16 Zejun Liu , Yunshan Chen , Chengxi Xie , Yugui Xie , Huan Liu

Emotion Recognition (ER) is the process of analyzing and identifying human emotions from sensing data. Currently, the field heavily relies on facial expression recognition (FER) because visual channel conveys rich emotional cues. However,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Kejun Liu , Yuanyuan Liu , Lin Wei , Chang Tang , Yibing Zhan , Zijing Chen , Zhe Chen

Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the…

声音 · 计算机科学 2022-07-19 Amir Shirian , Krishna Somandepalli , Victor Sanchez , Tanaya Guha

In recent years, there has been a significant increase in applications of multimodal signal processing and analysis, largely driven by the increased availability of multimodal datasets and the rapid progress in multimodal learning systems.…

图像与视频处理 · 电气工程与系统科学 2024-05-22 Hadi Hadizadeh , S. Faegheh Yeganli , Bahador Rashidi , Ivan V. Bajić

Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed…

机器学习 · 计算机科学 2022-05-03 Ahmed Abdou , Ekta Sood , Philipp Müller , Andreas Bulling

Emotion recognition is essential for applications in affective computing and behavioral prediction, but conventional systems relying on single-modality data often fail to capture the complexity of affective states. To address this…

多媒体 · 计算机科学 2025-09-08 Jianlu Wang , Yanan Wang , Tong Liu

In this paper, we address the task of utterance level emotion recognition in conversations using commonsense knowledge. We propose COSMIC, a new framework that incorporates different elements of commonsense such as mental states, events,…

计算与语言 · 计算机科学 2020-10-07 Deepanway Ghosal , Navonil Majumder , Alexander Gelbukh , Rada Mihalcea , Soujanya Poria

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

多媒体 · 计算机科学 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocentric video-based models capture rich semantic information and have demonstrated strong performance…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Baiyu Chen , Wilson Wongso , Zechen Li , Yonchanok Khaokaew , Hao Xue , Flora Salim

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text…

计算与语言 · 计算机科学 2026-02-27 Soumya Dutta , Smruthi Balaji , Sriram Ganapathy

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety. Current…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Ege Özsoy , Chantal Pellegrini , Tobias Czempiel , Felix Tristram , Kun Yuan , David Bani-Harouni , Ulrich Eck , Benjamin Busam , Matthias Keicher , Nassir Navab

Image-guided story ending generation (IgSEG) is to generate a story ending based on given story plots and ending image. Existing methods focus on cross-modal feature fusion but overlook reasoning and mining implicit information from story…

计算机视觉与模式识别 · 计算机科学 2023-01-30 Yucheng Zhou , Guodong Long

As the result of the growing importance of the Human Computer Interface system, understanding human's emotion states has become a consequential ability for the computer. This paper aims to improve the performance of emotion recognition by…

定量方法 · 定量生物学 2018-09-25 Kuan Tung , Po-Kang Liu , Yu-Chuan Chuang , Sheng-Hui Wang , An-Yeu Wu

Cross-modal generalization aims to learn a shared discrete representation space from multimodal pairs, enabling knowledge transfer across unannotated modalities. However, achieving a unified representation for all modality pairs requires…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Yan Xia , Hai Huang , Minghui Fang , Zhou Zhao

EEG is a non-invasive, safe, and low-risk method to record electrophysiological signals inside the brain. Especially with recent technology developments like dry electrodes, consumer-grade EEG devices, and rapid advances in machine…

机器学习 · 计算机科学 2025-06-23 Tri Duc Ly , Gia H. Ngo