中文
相关论文

相关论文: Recursive Joint Attention for Audio-Visual Fusion …

200 篇论文

Due to the severe lack of labeled data, existing methods of medical visual question answering usually rely on transfer learning to obtain effective image feature representation and use cross-modal fusion of visual and linguistic features to…

多媒体 · 计算机科学 2021-05-04 Haifan Gong , Guanqi Chen , Sishuo Liu , Yizhou Yu , Guanbin Li

Intelligently reasoning about the world often requires integrating data from multiple modalities, as any individual modality may contain unreliable or incomplete information. Prior work in multimodal learning fuses input modalities only…

机器学习 · 计算机科学 2020-11-17 George Barnum , Sabera Talukder , Yisong Yue

Music emotion recognition (MER), a sub-task of music information retrieval (MIR), has developed rapidly in recent years. However, the learning of affect-salient features remains a challenge. In this paper, we propose an end-to-end…

声音 · 计算机科学 2022-07-01 Zi Huang , Shulei Ji , Zhilan Hu , Chuangjian Cai , Jing Luo , Xinyu Yang

Empathetic and coherent responses are critical in auto-mated chatbot-facilitated psychotherapy. This study addresses the challenge of enhancing the emotional and contextual understanding of large language models (LLMs) in psychiatric…

计算与语言 · 计算机科学 2025-03-12 Abdur Rasool , Muhammad Irfan Shahzad , Hafsa Aslam , Vincent Chan , Muhammad Ali Arshad

We present Perceiver-VL, a vision-and-language framework that efficiently handles high-dimensional multimodal inputs such as long videos and text. Powered by the iterative latent cross-attention of Perceiver, our framework scales with…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Zineng Tang , Jaemin Cho , Jie Lei , Mohit Bansal

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality…

音频与语音处理 · 电气工程与系统科学 2025-04-01 Jinming Chen , Jingyi Fang , Yuanzhong Zheng , Yaoxuan Wang , Haojun Fei

This paper introduces a new multi-modal model based on the Transformer architecture and tensor product fusion strategy, combining BERT's text vectors and ViT's image vectors to classify students' psychological conditions, with an accuracy…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Ao Xiang , Zongqing Qi , Han Wang , Qin Yang , Danqing Ma

Composed Video Retrieval (CoVR) facilitates video retrieval by combining visual and textual queries. However, existing CoVR frameworks typically fuse multimodal inputs in a single stage, achieving only marginal gains over initial baseline.…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Yuqian Zheng , Mariana-Iuliana Georgescu

The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in this paper use the…

计算与语言 · 计算机科学 2023-06-13 Théo Deschamps-Berger , Lori Lamel , Laurence Devillers

Computer interfaces are advancing towards using multi-modalities to enable better human-computer interactions. The use of automatic emotion recognition (AER) can make the interactions natural and meaningful thereby enhancing the user…

声音 · 计算机科学 2025-03-26 Upasana Tiwari , Rupayan Chakraborty , Sunil Kumar Kopparapu

Continuous emotion recognition in terms of valence and arousal under in-the-wild (ITW) conditions remains a challenging problem due to large variations in appearance, head pose, illumination, occlusions, and subject-specific patterns of…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Elena Ryumina , Maxim Markitantov , Alexandr Axyonov , Dmitry Ryumin , Mikhail Dolgushin , Denis Dresvyanskiy , Alexey Karpov

We present a novel LSTM cell architecture capable of learning both intra- and inter-perspective relationships available in visual sequences captured from multiple perspectives. Our architecture adopts a novel recurrent joint learning…

计算机视觉与模式识别 · 计算机科学 2021-05-07 Alireza Sepas-Moghaddam , Fernando Pereira , Paulo Lobato Correia , Ali Etemad

Classifying group-level emotions is a challenging task due to complexity of video, in which not only visual, but also audio information should be taken into consideration. Existing works on multimodal emotion recognition are using bulky…

计算机视觉与模式识别 · 计算机科学 2021-11-12 Lev Evtodienko

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa

In this work, we explore the impact of visual modality in addition to speech and text for improving the accuracy of the emotion detection system. The traditional approaches tackle this task by fusing the knowledge from the various…

机器学习 · 计算机科学 2020-04-24 Seunghyun Yoon , Subhadeep Dey , Hwanhee Lee , Kyomin Jung

This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio segments to extract visual and audio features. Through the…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Jun Yu , Gongpeng Zhao , Yongqi Wang , Zhihong Wei , Yang Zheng , Zerui Zhang , Zhongpeng Cai , Guochen Xie , Jichao Zhu , Wangyuan Zhu

Multi-modal emotion recognition is challenging due to the difficulty of extracting features that capture subtle emotional differences. Understanding multi-modal interactions and connections is key to building effective bimodal speech…

声音 · 计算机科学 2025-03-25 Jiachen Luo , Huy Phan , Lin Wang , Joshua D. Reiss

In the latest social networks, more and more people prefer to express their emotions in videos through text, speech, and rich facial expressions. Multimodal video emotion analysis techniques can help understand users' inner world…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Qinglan Wei , Xuling Huang , Yuan Zhang

Spatial and temporal relationships, both short-range and long-range, between objects in videos, are key cues for recognizing actions. It is a challenging problem to model them jointly. In this paper, we first present a new variant of Long…

计算机视觉与模式识别 · 计算机科学 2020-04-28 Zexi Chen , Bharathkumar Ramachandra , Tianfu Wu , Ranga Raju Vatsavai