中文
相关论文

相关论文: Hierarchical Vision-Language Interaction for Facia…

200 篇论文

Due to its importance in facial behaviour analysis, facial action unit (AU) detection has attracted increasing attention from the research community. Leveraging the online knowledge distillation framework, we propose the ``FANTrans" method…

计算机视觉与模式识别 · 计算机科学 2022-11-14 Jing Yang , Jie Shen , Yiming Lin , Yordan Hristov , Maja Pantic

Automatic facial action unit (AU) recognition is a challenging task due to the scarcity of manual annotations. To alleviate this problem, a large amount of efforts has been dedicated to exploiting various methods which leverage numerous…

计算机视觉与模式识别 · 计算机科学 2021-08-02 Jingwei Yan , Jingjing Wang , Qiang Li , Chunmao Wang , Shiliang Pu

Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrates three subtasks: emotion recognition, facial Action Unit (AU) recognition, and AU-based…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Jiulong Wu , Yucheng Shen , Lingyong Yan , Haixin Sun , Deguo Xia , Jizhou Huang , Min Cao

Facial Action Unit (AU) detection has gained significant attention as it enables the breakdown of complex facial expressions into individual muscle movements. In this paper, we revisit two fundamental factors in AU detection: diverse and…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Mang Ning , Albert Ali Salah , Itir Onal Ertugrul

Since Facial Action Unit (AU) annotations require domain expertise, common AU datasets only contain a limited number of subjects. As a result, a crucial challenge for AU detection is addressing identity overfitting. We find that AUs and…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Zhipeng Hu , Wei Zhang , Lincheng Li , Yu Ding , Wei Chen , Zhigang Deng , Xin Yu

Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion…

计算机视觉与模式识别 · 计算机科学 2024-07-09 R. Gnana Praveen , Eric Granger , Patrick Cardinal

Facial Action Coding System is an important approach of facial expression analysis.This paper describes our submission to the third Affective Behavior Analysis (ABAW) 2022 competition. We proposed a transfomer based model to detect facial…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Lingfeng Wang , Shisen Wang , Jin Qi

Detecting action units (AUs) on human faces is challenging because various AUs make subtle facial appearance change over various regions at different scales. Current works have attempted to recognize AUs by emphasizing important regions.…

计算机视觉与模式识别 · 计算机科学 2019-08-27 Chen Ma , Li Chen , Junhai Yong

We propose a novel convolutional neural network approach to address the fine-grained recognition problem of multi-view dynamic facial action unit detection. We leverage recent gains in large-scale object recognition by formulating the task…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Andres Romero , Juan Leon , Pablo Arbelaez

Action Unit (AU) detection plays an important role for facial expression recognition. To the best of our knowledge, there is little research about AU analysis for micro-expressions. In this paper, we focus on AU detection in…

计算机视觉与模式识别 · 计算机科学 2020-04-13 Yante Li , Xiaohua Huang , Guoying Zhao

In vision and linguistics; the main input modalities are facial expressions, speech patterns, and the words uttered. The issue with analysis of any one mode of expression (Visual, Verbal or Vocal) is that lot of contextual information can…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Kunjal Panchal

We consider the problem of Visual Question Answering (VQA). Given an image and a free-form, open-ended, question, expressed in natural language, the goal of VQA system is to provide accurate answer to this question with respect to the…

计算机视觉与模式识别 · 计算机科学 2021-06-07 Tanzila Rahman , Shih-Han Chou , Leonid Sigal , Giuseppe Carenini

A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of semantic ambiguity, complex environmental layouts, and domain-specific terminology common…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Ping Li , Bartlomiej Brzozka

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

多媒体 · 计算机科学 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

In this work, we focus on leveraging facial cues beyond the lip region for robust Audio-Visual Speech Enhancement (AVSE). The facial region, encompassing the lip region, reflects additional speech-related attributes such as gender, skin…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Feixiang Wang , Shuang Yang , Shiguang Shan , Xilin Chen

Dynamic Facial Expression Recognition(DFER) is a rapidly evolving field of research that focuses on the recognition of time-series facial expressions. While previous research on DFER has concentrated on feature learning from a deep learning…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Feng Liu , Lingna Gu , Chen Shi , Xiaolan Fu

Recognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Current deep-learning-based algorithms may achieve impressive…

计算机视觉与模式识别 · 计算机科学 2021-04-05 Tao Pu , Tianshui Chen , Yuan Xie , Hefeng Wu , Liang Lin

The muscular activities caused the activation of certain AUs for every facial expression at the certain duration of time throughout the facial expression. This paper presents the methods to recognise facial Action Unit (AU) using facial…

计算机视觉与模式识别 · 计算机科学 2017-12-04 N. Hussain , H. Ujir , I. Hipiny , J-L Minoi

State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal (contrastive) or…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Amanpreet Singh , Ronghang Hu , Vedanuj Goswami , Guillaume Couairon , Wojciech Galuba , Marcus Rohrbach , Douwe Kiela

Effective human-agent interaction (HAI) relies on accurate and adaptive perception of human emotional states. While multimodal deep learning models - leveraging facial expressions, speech, and textual cues - offer high accuracy in emotion…

机器学习 · 计算机科学 2025-12-15 Matvey Nepomnyaschiy , Oleg Pereziabov , Anvar Tliamov , Stanislav Mikhailov , Ilya Afanasyev