一种语言,两种脚本:探究 LLM 概念表征中的脚本不变性
摘要
稀疏自编码器 (SAE) 学习的特征是表示抽象意义,还是绑定于文本书写方式?我们使用塞尔维亚文字双写作作为受控测试环境:塞尔维亚通用拉丁文和西里尔文两种书写方式可完美的字符映射互换,使我们能够在保持意义恒定的前提下变化文字编码。 Crucially, these scripts are tokenized completely differently, sharing no tokens whatsoever. Analyzing SAE feature activations across the Gemma model family (270M-27B parameters), we find that identical sentences in different Serbian scripts activate highly overlapping features, far exceeding random baselines. Strikingly, changing script causes less representational divergence than paraphrasing within the same script, suggesting SAE features prioritize meaning over orthographic form. Cross-script cross-paraphrase comparisons provide evidence against memorization, as these combinations rarely co-occur in training data yet still exhibit substantial feature overlap. This script invariance strengthens with model scale. Taken together, our findings suggest that SAE features can capture semantics at a level of abstraction above surface tokenization, and we propose Serbian digraphia as a general evaluation paradigm for probing the abstractness of learned representations.
引用
@article{arxiv.2603.08869,
title = {One Language, Two Scripts: Probing Script-Invariance in LLM Concept Representations},
author = {Sripad Karne},
journal= {arXiv preprint arXiv:2603.08869},
year = {2026}
}
备注
Accepted at the UCRL Workshop at ICLR 2026