English
Related papers

Related papers: Detail-Enhanced Intra- and Inter-modal Interaction…

200 papers

Video periocular recognition is the task of recognizing an individual's identity based on the region around an individual's eyes. The periocular area is one of the most discriminative regions of the human face, making it suitable for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Luiz G F Carreira , Breno A Mariano , Victor H C de Melo , David Menotti , William Robson Schwartz

Anomaly recognition plays a vital role in surveillance, transportation, healthcare, and public safety. However, most existing approaches rely solely on visual data, making them unreliable under challenging conditions such as occlusion, low…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Amjid Ali , Zulfiqar Ahmad Khan , Altaf Hussain , Muhammad Munsif , Adnan Hussain , Sung Wook Baik

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Ina Salaj , Arijit Biswas

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Autonomous driving technology has advanced significantly, yet detecting driving anomalies remains a major challenge due to the long-tailed distribution of driving events. Existing methods primarily rely on single-modal road condition video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Long Zhouxiang , Ovanes Petrosian

Generating music with emotion similar to that of an input video is a very relevant issue nowadays. Video content creators and automatic movie directors benefit from maintaining their viewers engaged, which can be facilitated by producing…

Sound · Computer Science 2020-04-07 Gwenaelle Cunha Sergio , Minho Lee

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

The purpose of emotion recognition in conversation (ERC) is to identify the emotion category of an utterance based on contextual information. Previous ERC methods relied on simple connections for cross-modal fusion and ignored the…

Computation and Language · Computer Science 2024-05-29 Haoxiang Shi , Xulong Zhang , Ning Cheng , Yong Zhang , Jun Yu , Jing Xiao , Jianzong Wang

The performance of CLIP in dynamic facial expression recognition (DFER) task doesn't yield exceptional results as observed in other CLIP-based classification tasks. While CLIP's primary objective is to achieve alignment between images and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Zeng Tao , Yan Wang , Junxiong Lin , Haoran Wang , Xinji Mai , Jiawen Yu , Xuan Tong , Ziheng Zhou , Shaoqi Yan , Qing Zhao , Liyuan Han , Wenqiang Zhang

2D+3D facial expression recognition (FER) can effectively cope with illumination changes and pose variations by simultaneously merging 2D texture and more robust 3D depth information. Most deep learning-based approaches employ the simple…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Mingzhe Sui , Hanting Li , Zhaoqing Zhu , Feng Zhao

Associating driver attention with driving scene across two fields of views (FOVs) is a hard cross-domain perception problem, which requires comprehensive consideration of cross-view mapping, dynamic driving scene analysis, and driver status…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Jun Zhou , Chunsheng Liu , Faliang Chang , Wenqian Wang , Penghui Hao , Yiming Huang , Zhiqiang Yang

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

Multimedia · Computer Science 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

This paper introduces our method for the Emotional Reaction Intensity (ERI) Estimation Challenge, in CVPR 2023: 5th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW). Based on the multimodal data provided by the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Shangfei Wang , Jiaqiang Wu , Feiyi Zheng , Xin Li , Xuewei Li , Suwen Wang , Yi Wu , Yanan Chang , Xiangyu Miao

Emotion recognition (ER) from speech signals is a robust approach since it cannot be imitated like facial expression or text based sentiment analysis. Valuable information underlying the emotions are significant for human-computer…

Sound · Computer Science 2023-12-19 David Hason Rudd , Huan Huo , Guandong Xu

Multi-modal emotion recognition is challenging due to the difficulty of extracting features that capture subtle emotional differences. Understanding multi-modal interactions and connections is key to building effective bimodal speech…

Sound · Computer Science 2025-03-25 Jiachen Luo , Huy Phan , Lin Wang , Joshua D. Reiss

This paper discusses the benefits of incorporating multimodal data for improving latent emotion recognition accuracy, focusing on micro-expression (ME) and physiological signals (PS). The proposed approach presents a novel multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Liangfei Zhang , Yifei Qian , Ognjen Arandjelovic , Anthony Zhu

In this paper, we present a multimodal approach to simultaneously analyze facial movements and several peripheral physiological signals to decode individualized affective experiences under positive and negative emotional contexts, while…

Computer Vision and Pattern Recognition · Computer Science 2018-11-20 Yu Yin , Mohsen Nabian , Miolin Fan , ChunAn Chou , Maria Gendron , Sarah Ostadabbas

In this paper, we propose a novel deep inductive transfer learning framework, named feature distribution adaptation network, to tackle the challenging multi-modal speech emotion recognition problem. Our method aims to use deep transfer…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Shaokai Li , Yixuan Ji , Peng Song , Haoqin Sun , Wenming Zheng

The goal of the audio-visual segmentation (AVS) task is to segment the sounding objects in the video frames using audio cues. However, current fusion-based methods have the performance limitations due to the small receptive field of…

Sound · Computer Science 2023-07-26 Jinxiang Liu , Chen Ju , Chaofan Ma , Yanfeng Wang , Yu Wang , Ya Zhang

Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and fine textures from visible imaging. Such accurate fusion is essential for real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zhenyu Sun , Luobin Zhang , Axi Niu , Haishen Wang , Qingsen Yan
‹ Prev 1 8 9 10 Next ›