English
Related papers

Related papers: AUD-TGN: Advancing Action Unit Detection with Temp…

200 papers

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

Computers and Society · Computer Science 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature…

Multimedia · Computer Science 2023-03-07 Zhongweiyang Xu , Xulin Fan , Mark Hasegawa-Johnson

Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Junhyeong Byeon , Jeongyeol Kim , Sejoon Lim

Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Wentao Zhu

Multimodal camera-LiDAR fusion technology has found extensive application in 3D object detection, demonstrating encouraging performance. However, existing methods exhibit significant performance degradation in challenging scenarios…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Sixian Liu , Chen Xu , Qiang Wang , Donghai Shi , Yiwen Li

Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Namho Kim , Junhwa Kim

Multimodal emotion recognition often suffers from performance degradation in valence-arousal estimation due to noise and misalignment between audio and visual modalities. To address this challenge, we introduce TAGF, a Time-aware Gated…

Multimedia · Computer Science 2025-07-04 Yubeen Lee , Sangeun Lee , Chaewon Park , Junyeop Cha , Eunil Park

We used two multimodal models for continuous valence-arousal recognition using visual, audio, and linguistic information. The first model is the same as we used in ABAW2 and ABAW3, which employs the leader-follower attention. The second…

Multimedia · Computer Science 2023-04-18 Su Zhang , Ziyuan Zhao , Cuntai Guan

In this work we tackle the task of video-based visual emotion recognition in the wild. Standard methodologies that rely solely on the extraction of bodily and facial features often fall short of accurate emotion prediction in cases where…

Computer Vision and Pattern Recognition · Computer Science 2022-02-03 Ioannis Pikoulis , Panagiotis P. Filntisis , Petros Maragos

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-27 Wim Boes , Hugo Van hamme

To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal sensors, visual and non-visual tend to provide complementary…

Computer Vision and Pattern Recognition · Computer Science 2022-11-09 Hyeongju Choi , Apoorva Beedu , Harish Haresamudram , Irfan Essa

Traditional approaches in speech emotion recognition, such as LSTM, CNN, RNN, SVM, and MLP, have limitations such as difficulty capturing long-term dependencies in sequential data, capturing the temporal dynamics, and struggling to capture…

Sound · Computer Science 2023-08-10 Samiul Islam , Md. Maksudul Haque , Abu Jobayer Md. Sadat

Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional…

Robotics · Computer Science 2024-09-19 Youssef Mohamed , Severin Lemaignan , Arzu Guneysu , Patric Jensfelt , Christian Smith

In this paper we propose a fusion approach to continuous emotion recognition that combines visual and auditory modalities in their representation spaces to predict the arousal and valence levels. The proposed approach employs a pre-trained…

Machine Learning · Computer Science 2019-06-26 Juan D. S. Ortega , Patrick Cardinal , Alessandro L. Koerich

In this paper, we examine the use of data from multiple sensing modes, i.e., accelerometry and global navigation satellite system (GNSS), for classifying animal behavior. We extract three new features from the GNSS data, namely, distance…

Machine Learning · Computer Science 2022-10-27 Reza Arablouei , Ziwei Wang , Greg J. Bishop-Hurley , Jiajun Liu

Facial Action Units (AUs) detection is a cornerstone of objective facial expression analysis and a critical focus in affective computing. Despite its importance, AU detection faces significant challenges, such as the high cost of AU…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Bohao Xing , Kaishen Yuan , Zitong Yu , Xin Liu , Heikki Kälviäinen

This paper addresses the question of emotion classification. The task consists in predicting emotion labels (taken among a set of possible labels) best describing the emotions contained in short video clips. Building on a standard framework…

Computer Vision and Pattern Recognition · Computer Science 2017-09-22 Valentin Vielzeuf , Stéphane Pateux , Frédéric Jurie

Expressions and facial action units (AUs) are two levels of facial behavior descriptors. Expression auxiliary information has been widely used to improve the AU detection performance. However, most existing expression representations can…

Computer Vision and Pattern Recognition · Computer Science 2022-10-31 Rudong An , Wei Zhang , Hao Zeng , Wei Chen , Zhigang Deng , Yu Ding