English

Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs

Computation and Language 2026-01-27 v1 Artificial Intelligence

Abstract

Large Language Models (LLMs) offer strong generative capabilities, but many applications require explicit and \textit{fine-grained} control over specific textual concepts, such as humor, persuasiveness, or formality. Prior approaches in prompting and representation engineering can provide coarse or single-attribute control, but systematic evaluation of multi-attribute settings remains limited. We introduce an evaluation framework for fine-grained controllability for both single- and dual-concept scenarios, focusing on linguistically distinct concept pairs (e.g., persuasiveness vs.~humor). Surprisingly, across multiple LLMs and generative tasks, we find that performance often drops in the dual-concept setting, even though the chosen concepts should in principle be separable. This reveals a fundamental limitation of naive prompting-based control: models struggle with compositionality even when concepts are intuitively independent. Our framework provides systematic evidence of this gap and offers a principled approach for measuring the ability of future methods for multi-concept control.

Keywords

Cite

@article{arxiv.2601.18483,
  title  = {Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs},
  author = {Arya Labroo and Ivaxi Sheth and Vyas Raina and Amaani Ahmed and Mario Fritz},
  journal= {arXiv preprint arXiv:2601.18483},
  year   = {2026}
}

Comments

Accepted for publication at EACL main conference

R2 v1 2026-07-01T09:20:25.838Z