English
Related papers

Related papers: A$^{3}$lign-DFER: Pioneering Comprehensive Dynamic…

200 papers

Recently, Contrastive Language-Image Pre-training (CLIP) has shown promising performance in domain-specific data (e.g., biology), and has attracted increasing research attention. Existing works generally focus on collecting extensive…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Junjie Wu , Jiangtao Xie , Zhaolin Zhang , Qilong Wang , Qinghua Hu , Peihua Li , Sen Xu

Facial expression recognition (FER) is a fundamental task in affective computing with applications in human-computer interaction, mental health analysis, and behavioral understanding. In this paper, we propose SMILE-VLM, a self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Muzammil Behzad

We propose DiffCLIP, a novel vision-language model that extends the differential attention mechanism to CLIP architectures. Differential attention was originally developed for large language models to amplify relevant context while…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Hasan Abed Al Kader Hammoud , Bernard Ghanem

Vision-language models (VLMs) like CLIP excel in zero-shot learning by aligning image and text representations through contrastive pretraining. Existing approaches to unsupervised adaptation (UA) for fine-grained classification with VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Eman Ali , Sathira Silva , Chetan Arora , Muhammad Haris Khan

This paper proposes a novel 4D Facial Expression Recognition (FER) method using Collaborative Cross-domain Dynamic Image Network (CCDN). Given a 4D data of face scans, we first compute its geometrical images, and then combine their…

Computer Vision and Pattern Recognition · Computer Science 2020-02-10 Muzammil Behzad , Nhat Vo , Xiaobai Li , Guoying Zhao

Compared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natural expression recognition scene. However, DFER is often…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Hanting Li , Hongjing Niu , Zhaoqing Zhu , Feng Zhao

Continual learning with vision-language models like CLIP offers a pathway toward scalable machine learning systems by leveraging its transferable representations. Existing CLIP-based methods adapt the pre-trained image encoder by adding…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Mao-Lin Luo , Zi-Hao Zhou , Tong Wei , Min-Ling Zhang

The increase of web-scale weakly labelled image-text pairs have greatly facilitated the development of large-scale vision-language models (e.g., CLIP), which have shown impressive generalization performance over a series of downstream…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Lianyu Hu , Tongkai Shi , Liqing Gao , Zekang Liu , Wei Feng

Facial expression plays an important role in understanding human emotions. Most recently, deep learning based methods have shown promising for facial expression recognition. However, the performance of the current state-of-the-art facial…

Computer Vision and Pattern Recognition · Computer Science 2021-12-10 Ping Liu , Yunchao Wei , Zibo Meng , Weihong Deng , Joey Tianyi Zhou , Yi Yang

Micro expression recognition (MER) is crucial for inferring genuine emotion. Applying a multimodal large language model (MLLM) to this task enables spatio-temporal analysis of facial motion and provides interpretable descriptions. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Ren Zhang , Huilai Li , Chao qi , Guoliang Xu , Tianyu Zhou , Wei wei , Jianqin Yin

An automatic Facial Expression Recognition (FER) model with Adaboost face detector, feature selection based on manifold learning and synergetic prototype based classifier has been proposed. Improved feature selection method and proposed…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Chendi Wang

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xinwei He , Yansong Zheng , Qianru Han , Zhichuan Wang , Yuxuan Cai , Yang Zhou , Jingbo Xia , Yulong Wang , Jinhai Xiang , Xiang Bai

Human affective behavior analysis aims to delve into human expressions and behaviors to deepen our understanding of human emotions. Basic expression categories (EXPR) and Action Units (AUs) are two essential components in this analysis,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Li Lin , Sarah Papabathini , Xin Wang , Shu Hu

Facial Expression Recognition (FER) in the wild is still challenging due to uncontrolled variations in pose, occlusion, and illumination. Most existing attention-based methods primarily rely on visual appearance cues, suffering from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Jiaxin Wang , Muwei Jian , Hui Yu , Junyu Dong , Yifan Xia

Contrastive Language-Image Pretraining (CLIP) achieves strong generalization in vision-language tasks by aligning images and texts in a shared embedding space. However, recent findings show that CLIP-like models still underutilize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Weiheng Zhao , Zilong Huang , Jiashi Feng , Xinggang Wang

Federated learning (FL) enables multiple clients to collaboratively train machine learning models without exposing local data, balancing performance and privacy. However, domain shift and label heterogeneity across clients often hinder the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Yubin Zheng , Pak-Hei Yeung , Jing Xia , Tianjie Ju , Peng Tang , Weidong Qiu , Jagath C. Rajapakse

In 2D+3D facial expression recognition (FER), existing methods generate multi-view geometry maps to enhance the depth feature representation. However, this may introduce false estimations due to local plane fitting from incomplete point…

Computer Vision and Pattern Recognition · Computer Science 2020-11-18 Yang Jiao , Yi Niu , Trac D. Tran , Guangming Shi

Despite the remarkable performance of vision language models (VLMs) such as Contrastive Language Image Pre-training (CLIP), the large size of these models is a considerable obstacle to their use in federated learning (FL) systems where the…

Machine Learning · Computer Science 2025-03-11 Yihang Wu , Ahmad Chaddad , Christian Desrosiers , Tareef Daqqaq , Reem Kateb

Dynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on a single snapshot. This temporal analysis provides richer…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Yin Chen , Jia Li , Yu Zhang , Zhenzhen Hu , Shiguang Shan , Meng Wang , Richang Hong

Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With the emergence of large-scale unlabeled human image datasets, there is an increasing need…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Mingshuang Luo , Ruibing Hou , Bo Chao , Hong Chang , Zimo Liu , Yaowei Wang , Shiguang Shan