English

How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders

Computation and Language 2025-10-13 v2 Machine Learning

Abstract

This study explores how bilingual language models develop complex internal representations. We employ sparse autoencoders to analyze internal representations of bilingual language models with a focus on the effects of training steps, layers, and model sizes. Our analysis shows that language models first learn languages separately, and then gradually form bilingual alignments, particularly in the mid layers. We also found that this bilingual tendency is stronger in larger models. Building on these findings, we demonstrate the critical role of bilingual representations in model performance by employing a novel method that integrates decomposed representations from a fully trained model into a mid-training model. Our results provide insights into how language models acquire bilingual capabilities.

Keywords

Cite

@article{arxiv.2503.06394,
  title  = {How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders},
  author = {Tatsuro Inaba and Go Kamoda and Kentaro Inui and Masaru Isonuma and Yusuke Miyao and Yohei Oseki and Benjamin Heinzerling and Yu Takagi},
  journal= {arXiv preprint arXiv:2503.06394},
  year   = {2025}
}

Comments

13 pages, 17 figures, accepted to EMNLP 2025 findings

R2 v1 2026-06-28T22:12:30.242Z