Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners
Abstract
Consider a model trained at a single hospital to predict patient recovery, where the measured feature bundles the patient's true health signal () with a systematic artefact from that hospital's equipment (). Within that hospital, the artefact correlates with outcomes through unmeasured confounders such as patient demographics; an in-context learner rationally routes predictions through , not , and fails silently when deployed at a new hospital with different equipment. We formalise this as \emph{spurious routing in composite representations}: when a feature encodes a causal signal and a spurious signal in distinct subspaces, the ICL cannot determine which drives predictions. We prove that under ridge ICL, a linear in-context learner, this routing is unavoidable regardless of context size; TabPFN, a state-of-the-art pretrained tabular ICL model, shows qualitatively consistent behaviour empirically. We derive a closed-form characterisation, , confirmed at for linear ICL and for TabPFN. Contrary to intuition, larger context sharpens commitment to the dominant in-context signal, amplifying spurious routing by up to ; in the high-spurious corner, more expressive models show greater vulnerability empirically ( CSR gap at high entanglement). We introduce two lightweight mitigations: environment-stratified context construction and S-swap augmentation, that require only weak environment labels and no knowledge of the causal partition. S-swap reduces spurious routing by for linear ICL and for TabPFN, with TabPFN's causal sensitivity increasing simultaneously: the model does not become agnostic, it reroutes through the causal signal.
Cite
@article{arxiv.2607.25532,
title = {Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners},
author = {Athanasios Vlontzos and Giorgos Papanastasiou and Bernhard Kainz and Sotirios Tsaftaris},
journal= {arXiv preprint arXiv:2607.25532},
year = {2026}
}