中文
相关论文

相关论文: A Unified Transformer-based Network for multimodal…

200 篇论文

Equipping humanoid robots with the capability to understand emotional states of human interactants and express emotions appropriately according to situations is essential for affective human-robot interaction. However, enabling current…

机器人学 · 计算机科学 2026-03-18 Peizhen Li , Longbing Cao , Xiao-Ming Wu , Xiaohan Yu , Runze Yang

Automated emotion detection in speech is a challenging task due to the complex interdependence between words and the manner in which they are spoken. It is made more difficult by the available datasets; their small size and incompatible…

音频与语音处理 · 电气工程与系统科学 2020-11-16 Amith Ananthram , Kailash Karthik Saravanakumar , Jessica Huynh , Homayoon Beigi

Automatic emotion recognition has become increasingly important with the rise of AI, especially in fields like healthcare, education, and automotive systems. However, there is a lack of multimodal datasets, particularly involving body…

人工智能 · 计算机科学 2025-09-09 Seyed Muhammad Hossein Mousavi , Atiye Ilanloo

Emotion recognition from physiological signals remains challenging due to their non-stationary, noisy, and subject-dependent characteristics. This work presents, to the best of our knowledge, the first comprehensive application of liquid…

信号处理 · 电气工程与系统科学 2026-02-10 Anindya Bhattacharjee , Nittya Ananda Biswas , K. A. Shahriar , Adib Rahman

Emotion recognition in conversations (ERC) aims to predict the emotional state of each utterance by using multiple input types, such as text and audio. While Transformer-based models have shown strong performance in this task, they often…

音频与语音处理 · 电气工程与系统科学 2025-08-13 Zhining He , Yang Xiao

Biomedical signal processing extract meaningful information from physiological signals like electrocardiograms (ECGs), electroencephalograms (EEGs), and electromyograms (EMGs) to diagnose, monitor, and treat medical conditions and diseases…

信号处理 · 电气工程与系统科学 2025-08-13 Justin London

The flow-based generative model is a deep learning generative model, which obtains the ability to generate data by explicitly learning the data distribution. Theoretically its ability to restore data is stronger than other generative…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Gao Xu , Yuanpeng Long , Siwei Liu , Lijia Yang , Shimei Xu , Xiaoming Yao , Kunxian Shu

Multimodal sentiment analysis, a pivotal task in affective computing, seeks to understand human emotions by integrating cues from language, audio, and visual signals. While many recent approaches leverage complex attention mechanisms and…

计算与语言 · 计算机科学 2025-05-09 Nischal Mandal , Yang Li

There has been an encouraging progress in the affective states recognition models based on the single-modality signals as electroencephalogram (EEG) signals or peripheral physiological signals in recent years. However, multimodal…

信号处理 · 电气工程与系统科学 2023-06-02 Yuxuan Zhao , Xinyan Cao , Jinlong Lin , Dunshan Yu , Xixin Cao

Multimodal Emotion Recognition in Conversation (ERC) plays an influential role in the field of human-computer interaction and conversational robotics since it can motivate machines to provide empathetic services. Multimodal data modeling is…

多媒体 · 计算机科学 2023-11-23 Jiang Li , Xiaoping Wang , Guoqing Lv , Zhigang Zeng

In-vehicle emotion recognition underpins adaptive driver-assistance systems and, ultimately, occupant safety. However, practical deployment is hindered by (i) modality fragility - poor lighting and occlusions degrade vision-based methods;…

机器学习 · 计算机科学 2025-07-23 Baran Can Gül , Suraksha Nadig , Stefanos Tziampazis , Nasser Jazdi , Michael Weyrich

The field of affective computing has seen significant advancements in exploring the relationship between emotions and emerging technologies. This paper presents a novel and valuable contribution to this field with the introduction of a…

人工智能 · 计算机科学 2025-01-15 Nessrine Farhat , Amine Bohi , Leila Ben Letaifa , Rim Slama

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Nikita Drobyshev , Antoni Bigata Casademunt , Konstantinos Vougioukas , Zoe Landgraf , Stavros Petridis , Maja Pantic

Vision transformer (ViT) has been widely applied in many areas due to its self-attention mechanism that help obtain the global receptive field since the first layer. It even achieves surprising performance exceeding CNN in some vision…

计算机视觉与模式识别 · 计算机科学 2021-09-28 Hanting Li , Mingzhe Sui , Zhaoqing Zhu , Feng Zhao

Facial expression recognition has been an active research area over the past few decades, and it is still challenging due to the high intra-class variation. Traditional approaches for this problem rely on hand-crafted features such as SIFT,…

计算机视觉与模式识别 · 计算机科学 2019-02-05 Shervin Minaee , Amirali Abdolrashidi

Emotion recognition is significantly enhanced by integrating multimodal biosignals and IMU data from multiple domains. In this paper, we introduce a novel multi-scale attention-based LSTM architecture, combined with Squeeze-and-Excitation…

信号处理 · 电气工程与系统科学 2024-12-04 Pubudu L. Indrasiri , Bipasha Kashyap , Chandima Kolambahewage , Bahareh Nakisa , Kiran Ijaz , Pubudu N. Pathirana

We present a glasses type wearable device to detect emotions from a human face in an unobtrusive manner. The device is designed to gather multi channel responses from the user face naturally and continuously while the user is wearing it.…

人机交互 · 计算机科学 2024-10-30 Jangho Kwon , Laehyun Kim

Facial expression recognition is a topic of great interest in most fields from artificial intelligence and gaming to marketing and healthcare. The goal of this paper is to classify images of human faces into one of seven basic emotions. A…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Akash Saravanan , Gurudutt Perichetla , K. S. Gayathri

In this paper we propose a new approach for classifying the global emotion of images containing groups of people. To achieve this task, we consider two different and complementary sources of information: i) a global representation of the…

计算机视觉与模式识别 · 计算机科学 2018-07-11 Aarush Gupta , Dakshit Agrawal , Hardik Chauhan , Jose Dolz , Marco Pedersoli

In this paper, we explore the possibility of building a unified foundation model that can be adapted to both vision-only and text-only tasks. Starting from BERT and ViT, we design a unified transformer consisting of modality-specific…

计算机视觉与模式识别 · 计算机科学 2021-12-15 Qing Li , Boqing Gong , Yin Cui , Dan Kondratyuk , Xianzhi Du , Ming-Hsuan Yang , Matthew Brown