English

Is In-Context Universality Enough? MLPs are Also Universal In-Context

Machine Learning 2025-02-06 v1 Machine Learning Numerical Analysis Neural and Evolutionary Computing Numerical Analysis Probability

Abstract

The success of transformers is often linked to their ability to perform in-context learning. Recent work shows that transformers are universal in context, capable of approximating any real-valued continuous function of a context (a probability measure over XRd\mathcal{X}\subseteq \mathbb{R}^d) and a query xXx\in \mathcal{X}. This raises the question: Does in-context universality explain their advantage over classical models? We answer this in the negative by proving that MLPs with trainable activation functions are also universal in-context. This suggests the transformer's success is likely due to other factors like inductive bias or training stability.

Cite

@article{arxiv.2502.03327,
  title  = {Is In-Context Universality Enough? MLPs are Also Universal In-Context},
  author = {Anastasis Kratsios and Takashi Furuya},
  journal= {arXiv preprint arXiv:2502.03327},
  year   = {2025}
}
R2 v1 2026-06-28T21:33:40.870Z