English

Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

Computation and Language 2025-06-19 v1 Artificial Intelligence

Abstract

In generative commonsense reasoning tasks such as CommonGen, generative large language models (LLMs) compose sentences that include all given concepts. However, when focusing on instruction-following capabilities, if a prompt specifies a concept order, LLMs must generate sentences that adhere to the specified order. To address this, we propose Ordered CommonGen, a benchmark designed to evaluate the compositional generalization and instruction-following abilities of LLMs. This benchmark measures ordered coverage to assess whether concepts are generated in the specified order, enabling a simultaneous evaluation of both abilities. We conducted a comprehensive analysis using 36 LLMs and found that, while LLMs generally understand the intent of instructions, biases toward specific concept order patterns often lead to low-diversity outputs or identical results even when the concept order is altered. Moreover, even the most instruction-compliant LLM achieved only about 75% ordered coverage, highlighting the need for improvements in both instruction-following and compositional generalization capabilities.

Keywords

Cite

@article{arxiv.2506.15629,
  title  = {Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability},
  author = {Yusuke Sakai and Hidetaka Kamigaito and Taro Watanabe},
  journal= {arXiv preprint arXiv:2506.15629},
  year   = {2025}
}

Comments

ACL 2025 Main

R2 v1 2026-07-01T03:23:56.118Z