English

In-Context Learning of Linear Systems: Generalization Theory and Applications to Operator Learning

Machine Learning 2025-05-27 v3 Numerical Analysis Numerical Analysis Machine Learning

Abstract

We study theoretical guarantees for solving linear systems in-context using a linear transformer architecture. For in-domain generalization, we provide neural scaling laws that bound the generalization error in terms of the number of tasks and sizes of samples used in training and inference. For out-of-domain generalization, we find that the behavior of trained transformers under task distribution shifts depends crucially on the distribution of the tasks seen during training. We introduce a novel notion of task diversity and show that it defines a necessary and sufficient condition for pre-trained transformers generalize under task distribution shifts. We also explore applications of learning linear systems in-context, such as to in-context operator learning for PDEs. Finally, we provide some numerical experiments to validate the established theory.

Keywords

Cite

@article{arxiv.2409.12293,
  title  = {In-Context Learning of Linear Systems: Generalization Theory and Applications to Operator Learning},
  author = {Frank Cole and Yulong Lu and Wuzhe Xu and Tianhao Zhang},
  journal= {arXiv preprint arXiv:2409.12293},
  year   = {2025}
}

Comments

Includes new results for in-domain generalization and operator learning, code available at https://github.com/LuGroupUMN/ICL_Linear_Systems

R2 v1 2026-06-28T18:49:32.885Z