English

Human vs. LMMs: Exploring the Discrepancy in Emoji Interpretation and Usage in Digital Communication

Computer Vision and Pattern Recognition 2024-04-16 v2

Abstract

Leveraging Large Multimodal Models (LMMs) to simulate human behaviors when processing multimodal information, especially in the context of social media, has garnered immense interest due to its broad potential and far-reaching implications. Emojis, as one of the most unique aspects of digital communication, are pivotal in enriching and often clarifying the emotional and tonal dimensions. Yet, there is a notable gap in understanding how these advanced models, such as GPT-4V, interpret and employ emojis in the nuanced context of online interaction. This study intends to bridge this gap by examining the behavior of GPT-4V in replicating human-like use of emojis. The findings reveal a discernible discrepancy between human and GPT-4V behaviors, likely due to the subjective nature of human interpretation and the limitations of GPT-4V's English-centric training, suggesting cultural biases and inadequate representation of non-English cultures.

Keywords

Cite

@article{arxiv.2401.08212,
  title  = {Human vs. LMMs: Exploring the Discrepancy in Emoji Interpretation and Usage in Digital Communication},
  author = {Hanjia Lyu and Weihong Qi and Zhongyu Wei and Jiebo Luo},
  journal= {arXiv preprint arXiv:2401.08212},
  year   = {2024}
}

Comments

Accepted for publication in ICWSM 2024