English

Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts

Computation and Language 2025-05-21 v2 Artificial Intelligence

Abstract

Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM's concept representations can be captured as steering vectors (SVs), enabling the control of LLM behavior (e.g., towards generating harmful content). Our work takes a novel approach by exploring the intricate relationships between concept representations across different LLMs, drawing an intriguing parallel to Plato's Allegory of the Cave. In particular, we introduce a linear transformation method to bridge these representations and present three key findings: 1) Concept representations across different LLMs can be effectively aligned using simple linear transformations, enabling efficient cross-model transfer and behavioral control via SVs. 2) This linear transformation generalizes across concepts, facilitating alignment and control of SVs representing different concepts across LLMs. 3) A weak-to-strong transferability exists between LLM concept representations, whereby SVs extracted from smaller LLMs can effectively control the behavior of larger LLMs.

Keywords

Cite

@article{arxiv.2501.02009,
  title  = {Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts},
  author = {Youcheng Huang and Chen Huang and Duanyu Feng and Wenqiang Lei and Jiancheng Lv},
  journal= {arXiv preprint arXiv:2501.02009},
  year   = {2025}
}

Comments

ACL 2025 Main Camera Ready