English

To See or To Read: User Behavior Reasoning in Multimodal LLMs

Artificial Intelligence 2025-11-07 v1 Machine Learning

Abstract

Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost.

Keywords

Cite

@article{arxiv.2511.03845,
  title  = {To See or To Read: User Behavior Reasoning in Multimodal LLMs},
  author = {Tianning Dong and Luyi Ma and Varun Vasudevan and Jason Cho and Sushant Kumar and Kannan Achan},
  journal= {arXiv preprint arXiv:2511.03845},
  year   = {2025}
}

Comments

Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Efficient Reasoning

R2 v1 2026-07-01T07:23:33.699Z