English

Kiss up, Kick down: Exploring Behavioral Changes in Multi-modal Large Language Models with Assigned Visual Personas

Computation and Language 2024-10-07 v1

Abstract

This study is the first to explore whether multi-modal large language models (LLMs) can align their behaviors with visual personas, addressing a significant gap in the literature that predominantly focuses on text-based personas. We developed a novel dataset of 5K fictional avatar images for assignment as visual personas to LLMs, and analyzed their negotiation behaviors based on the visual traits depicted in these images, with a particular focus on aggressiveness. The results indicate that LLMs assess the aggressiveness of images in a manner similar to humans and output more aggressive negotiation behaviors when prompted with an aggressive visual persona. Interestingly, the LLM exhibited more aggressive negotiation behaviors when the opponent's image appeared less aggressive than their own, and less aggressive behaviors when the opponents image appeared more aggressive.

Keywords

Cite

@article{arxiv.2410.03181,
  title  = {Kiss up, Kick down: Exploring Behavioral Changes in Multi-modal Large Language Models with Assigned Visual Personas},
  author = {Seungjong Sun and Eungu Lee and Seo Yeon Baek and Seunghyun Hwang and Wonbyung Lee and Dongyan Nan and Bernard J. Jansen and Jang Hyun Kim},
  journal= {arXiv preprint arXiv:2410.03181},
  year   = {2024}
}

Comments

EMNLP 2024

R2 v1 2026-06-28T19:08:09.523Z