中文
相关论文

相关论文: GPT-4V with Emotion: A Zero-shot Benchmark for Gen…

200 篇论文

Emotion Recognition (ER) is the process of identifying human emotions from given data. Currently, the field heavily relies on facial expression recognition (FER) because facial expressions contain rich emotional cues. However, it is…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Yuanyuan Liu , Lin Wei , Kejun Liu , Yibing Zhan , Zijing Chen , Zhe Chen , Shiguang Shan

Emotion Recognition (ER) is the process of analyzing and identifying human emotions from sensing data. Currently, the field heavily relies on facial expression recognition (FER) because visual channel conveys rich emotional cues. However,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Kejun Liu , Yuanyuan Liu , Lin Wei , Chang Tang , Yibing Zhan , Zijing Chen , Zhe Chen

We present a novel generalized zero-shot algorithm to recognize perceived emotions from gestures. Our task is to map gestures to novel emotion categories not encountered in training. We introduce an adversarial, autoencoder-based…

计算机视觉与模式识别 · 计算机科学 2021-12-03 Abhishek Banerjee , Uttaran Bhattacharya , Aniket Bera

As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study comprehensively evaluates the language, vision, speech, and…

OpenAI's latest large vision-language model (LVLM), GPT-4V(ision), has piqued considerable interest for its potential in medical applications. Despite its promise, recent studies and internal reviews highlight its underperformance in…

计算与语言 · 计算机科学 2023-12-13 Pengcheng Chen , Ziyan Huang , Zhongying Deng , Tianbin Li , Yanzhou Su , Haoyu Wang , Jin Ye , Yu Qiao , Junjun He

This work investigates the capabilities of large language models (LLMs) in detecting and understanding human emotions through text. Drawing upon emotion models from psychology, we adopt an interdisciplinary perspective that integrates…

计算与语言 · 计算机科学 2025-03-10 Florian Lecourt , Madalina Croitoru , Konstantin Todorov

Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to inconsistencies annotators often have different preferences…

人工智能 · 计算机科学 2025-06-02 Xin Jing , Jiadong Wang , Iosif Tsangko , Andreas Triantafyllopoulos , Björn W. Schuller

Large language models have seen widespread adoption in math problem-solving. However, in geometry problems that usually require visual aids for better understanding, even the most advanced multi-modal models currently still face challenges…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Shihao Cai , Keqin Bao , Hangyu Guo , Jizhi Zhang , Jun Song , Bo Zheng

Human emotion recognition is an active research area in artificial intelligence and has made substantial progress over the past few years. Many recent works mainly focus on facial regions to infer human affection, while the surrounding…

计算机视觉与模式识别 · 计算机科学 2021-11-09 Nhat Le , Khanh Nguyen , Anh Nguyen , Bac Le

In machine learning, generalization against distribution shifts -- where deployment conditions diverge from the training scenarios -- is crucial, particularly in fields like climate modeling, biomedicine, and autonomous driving. The…

机器学习 · 计算机科学 2024-02-27 Zhongyi Han , Guanglin Zhou , Rundong He , Jindong Wang , Tailin Wu , Yilong Yin , Salman Khan , Lina Yao , Tongliang Liu , Kun Zhang

Large Language Models (LLM) have recently been shown to perform well at various tasks from language understanding, reasoning, storytelling, and information search to theory of mind. In an extension of this work, we explore the ability of…

计算与语言 · 计算机科学 2023-10-31 Nutchanon Yongsatianchot , Tobias Thejll-Madsen , Stacy Marsella

The project leverages advanced machine and deep learning techniques to address the challenge of emotion recognition by focusing on non-facial cues, specifically hands, body gestures, and gestures. Traditional emotion recognition systems…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Haoyang Liu

The ability to handle various emotion labels without dedicated training is crucial for building adaptable Emotion Recognition (ER) systems. Conventional ER models rely on training using fixed label sets and struggle to generalize beyond…

计算与语言 · 计算机科学 2025-05-26 Minxue Niu , Emily Mower Provost

Large language models, in particular generative pre-trained transformers (GPTs), show impressive results on a wide variety of language-related tasks. In this paper, we explore ChatGPT's zero-shot ability to perform affective computing tasks…

计算与语言 · 计算机科学 2023-09-06 Joost Broekens , Bernhard Hilpert , Suzan Verberne , Kim Baraka , Patrick Gebhard , Aske Plaat

Deep convolutional neural networks have been shown to successfully recognize facial emotions for the past years in the realm of computer vision. However, the existing detection approaches are not always reliable or explainable, we here…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Jiawen Wang , Leah Kawka

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to…

Emotion recognition is relevant for human behaviour understanding, where facial expression and speech recognition have been widely explored by the computer vision community. Literature in the field of behavioural psychology indicates that…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Maria Luísa Lima , Willams de Lima Costa , Estefania Talavera Martinez , Veronica Teichrieb

Open-Vocabulary Multimodal Emotion Recognition (OV-MER) aims to predict emotions without being constrained by label spaces, enabling fine-grained emotion understanding. Unlike traditional discriminative methods, OV-MER leverages generative…

人机交互 · 计算机科学 2026-02-10 Zheng Lian , Fan Zhang , Yazhou Zhang , Jianhua Tao , Rui Liu , Haoyu Chen , Xiaobai Li , Bin He

This paper extends recent investigations on the emotional reasoning abilities of Large Language Models (LLMs). Current research on LLMs has not directly evaluated the distinction between how LLMs predict the self-attribution of emotions and…

人工智能 · 计算机科学 2024-08-27 Ala N. Tak , Jonathan Gratch

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

计算与语言 · 计算机科学 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi