中文
相关论文

相关论文: FineCLIPER: Multi-modal Fine-grained CLIP for Dyna…

200 篇论文

Deriving an effective facial expression recognition component is important for a successful human-computer interaction system. Nonetheless, recognizing facial expression remains a challenging task. This paper describes a novel approach…

计算机视觉与模式识别 · 计算机科学 2019-04-15 Mundher Al-Shabi , Wooi Ping Cheah , Tee Connie

Automatic Facial Expression Recognition (FER) has attracted increasing attention in the last 20 years since facial expressions play a central role in human communication. Most FER methodologies utilize Deep Neural Networks (DNNs) that are…

计算机视觉与模式识别 · 计算机科学 2022-05-10 Andreas Psaroudakis , Dimitrios Kollias

Existing pedestrian attribute recognition (PAR) algorithms are mainly developed based on a static image. However, the performance is not reliable for images with challenging factors, such as heavy occlusion, motion blur, etc. In this work,…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Jun Zhu , Jiandong Jin , Zihan Yang , Xiaohao Wu , Xiao Wang

In this paper, we present SAFER, a novel system for emotion recognition from facial expressions. It employs state-of-the-art deep learning techniques to extract various features from facial images and incorporates contextual information,…

计算机视觉与模式识别 · 计算机科学 2023-06-19 Mijanur Palash , Bharat Bhargava

Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in…

多媒体 · 计算机科学 2023-05-11 Juan Hu , Xin Liao , Difei Gao , Satoshi Tsutsui , Qian Wang , Zheng Qin , Mike Zheng Shou

Deepfake techniques have been widely used for malicious purposes, prompting extensive research interest in developing Deepfake detection methods. Deepfake manipulations typically involve tampering with facial parts, which can result in…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Juan Hu , Xin Liao , Difei Gao , Satoshi Tsutsui , Qian Wang , Zheng Qin , Mike Zheng Shou

Multimodal fake news detection has attracted many research interests in social forensics. Many existing approaches introduce tailored attention mechanisms to guide the fusion of unimodal features. However, how the similarity of these…

计算机视觉与模式识别 · 计算机科学 2022-05-31 Yangming Zhou , Qichao Ying , Zhenxing Qian , Sheng Li , Xinpeng Zhang

Despite the remarkable performance of vision language models (VLMs) such as Contrastive Language Image Pre-training (CLIP), the large size of these models is a considerable obstacle to their use in federated learning (FL) systems where the…

机器学习 · 计算机科学 2025-03-11 Yihang Wu , Ahmad Chaddad , Christian Desrosiers , Tareef Daqqaq , Reem Kateb

Recently, facial expression recognition (FER) in the wild has gained a lot of researchers' attention because it is a valuable topic to enable the FER techniques to move from the laboratory to the real applications. In this paper, we focus…

计算机视觉与模式识别 · 计算机科学 2020-08-14 Xingxun Jiang , Yuan Zong , Wenming Zheng , Chuangao Tang , Wanchuang Xia , Cheng Lu , Jiateng Liu

Dynamic facial expression recognition has many useful applications in social networks, multimedia content analysis, security systems and others. This challenging process must be done under recurrent problems of image illumination and low…

计算机视觉与模式识别 · 计算机科学 2020-12-29 Ali Raza Shahid , Sheheryar Khan , Hong Yan

Facial expression captioning has found widespread application across various domains. Recently, the emergence of video Multimodal Large Language Models (MLLMs) has shown promise in general video understanding tasks. However, describing…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Jiaxing Zhao , Boyuan Sun , Xiang Chen , Xihan Wei

Facial expression recognition (FER) has emerged as an important component of human-computer interaction systems. Despite recent advancements in FER, performance often drops significantly for non-frontal facial images. We propose Contrastive…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Shuvendu Roy , Ali Etemad

Recent approaches have shown that large-scale vision-language models such as CLIP can improve semantic segmentation performance. These methods typically aim for pixel-level vision-language alignment, but often rely on low resolution image…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Anurag Das , Xinting Hu , Li Jiang , Bernt Schiele

While vision-language models like CLIP have advanced zero-shot surgical phase recognition, they struggle with fine-grained surgical activities, especially action triplets. This limitation arises because current CLIP formulations rely on…

计算机视觉与模式识别 · 计算机科学 2025-03-30 Saurav Sharma , Didier Mutter , Nicolas Padoy

TIReID aims to retrieve the image corresponding to the given text query from a pool of candidate images. Existing methods employ prior knowledge from single-modality pre-training to facilitate learning, but lack multi-modal correspondences.…

计算机视觉与模式识别 · 计算机科学 2022-10-20 Shuanglin Yan , Neng Dong , Liyan Zhang , Jinhui Tang

Continual learning with vision-language models like CLIP offers a pathway toward scalable machine learning systems by leveraging its transferable representations. Existing CLIP-based methods adapt the pre-trained image encoder by adding…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Mao-Lin Luo , Zi-Hao Zhou , Tong Wei , Min-Ling Zhang

In this paper, we propose an approach for Facial Expressions Recognition (FER) based on a deep multi-facial patches aggregation network. Deep features are learned from facial patches using deep sub-networks and aggregated within one deep…

计算机视觉与模式识别 · 计算机科学 2020-03-17 Ahmed Rachid Hazourli , Amine Djeghri , Hanan Salam , Alice Othmani

Dynamic facial expression recognition (DFER) is a task that estimates emotions from facial expression video sequences. For practical applications, accurately recognizing ambiguous facial expressions -- frequently encountered in in-the-wild…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Ryosuke Kawamura , Hideaki Hayashi , Shunsuke Otake , Noriko Takemura , Hajime Nagahara

Automated facial expression analysis has a variety of applications in human-computer interaction. Traditional methods mainly analyze prototypical facial expressions of no more than eight discrete emotions as a classification task. However,…

计算机视觉与模式识别 · 计算机科学 2018-05-04 Feng Zhou , Shu Kong , Charless Fowlkes , Tao Chen , Baiying Lei

Meaningful facial parts can convey key cues for both facial action unit detection and expression prediction. Textured 3D face scan can provide both detailed 3D geometric shape and 2D texture appearance cues of the face which are beneficial…

计算机视觉与模式识别 · 计算机科学 2018-03-16 Asim Jan , Huaxiong Ding , Hongying Meng , Liming Chen , Huibin Li