中文

大语言模型能否捕捉视频游戏的参与度?

计算机视觉与模式识别 2026-01-30 v2 人工智能 计算与语言 人机交互

摘要

Can out-of-the-box pretrained Large Language Models (LLMs) detect human affect successfully when observing a video? To address this question, for the first time, we evaluate comprehensively the capacity of popular LLMs for successfully predicting continuous affect annotations of videos when prompted by a sequence of text and video frames in a multimodal fashion. In this paper, we test LLMs' ability to correctly label changes of in-game engagement in 80 minutes of annotated videogame footage from 20 first-person shooter games of the GameVibe corpus. We run over 4,800 experiments to investigate the impact of LLM architecture, model size, input modality, prompting strategy, and ground truth processing method on engagement prediction. Our findings suggest that while LLMs rightfully claim human-like performance across multiple domains and able to outperform traditional machine learning baselines, they generally fall behind continuous experience annotations provided by humans. We examine some of the underlying causes for a fluctuating performance across games, highlight the cases where LLMs exceed expectations, and draw a roadmap for the further exploration of automated emotion labelling via LLMs.

关键词

引用

@article{arxiv.2502.04379,
  title  = {Can Large Language Models Capture Video Game Engagement?},
  author = {David Melhart and Matthew Barthet and Georgios N. Yannakakis},
  journal= {arXiv preprint arXiv:2502.04379},
  year   = {2026}
}

备注

This work has been submitted to the IEEE for publication