English

SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services

Computation and Language 2025-12-16 v2

Abstract

With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and platform intelligence. Existing benchmarks primarily focus on text-centric tasks, lacking coverage of the multimodal contexts prevalent in modern SNS ecosystems. In this paper, we introduce SNS-Bench-VL, a comprehensive multimodal benchmark designed to assess the performance of Vision-Language LLMs in real-world social media scenarios. SNS-Bench-VL incorporates images and text across 8 multimodal tasks, including note comprehension, user engagement analysis, information retrieval, and personalized recommendation. It comprises 4,001 carefully curated multimodal question-answer pairs, covering single-choice, multiple-choice, and open-ended tasks. We evaluate over 25 state-of-the-art multimodal LLMs, analyzing their performance across tasks. Our findings highlight persistent challenges in multimodal social context comprehension. We hope SNS-Bench-VL will inspire future research towards robust, context-aware, and human-aligned multimodal intelligence for next-generation social networking services.

Keywords

Cite

@article{arxiv.2505.23065,
  title  = {SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services},
  author = {Hongcheng Guo and Zheyong Xie and Shaosheng Cao and Boyang Wang and Weiting Liu and Anjie Le and Lei Li and Zhoujun Li},
  journal= {arXiv preprint arXiv:2505.23065},
  year   = {2025}
}

Comments

We found problems in the code while rechecking our implementation. These issues led to noticeable numerical discrepancies, making some of the reported results and conclusions potentially unreliable. Therefore, we request to withdraw this submission

R2 v1 2026-07-01T02:47:44.723Z