English

Exploring the Robustness of Language Models for Tabular Question Answering via Attention Analysis

Computation and Language 2025-08-27 v4 Artificial Intelligence

Abstract

Large Language Models (LLMs), already shown to ace various unstructured text comprehension tasks, have also remarkably been shown to tackle table (structured) comprehension tasks without specific training. Building on earlier studies of LLMs for tabular tasks, we probe how in-context learning (ICL), model scale, instruction tuning, and domain bias affect Tabular QA (TQA) robustness by testing LLMs, under diverse augmentations and perturbations, on diverse domains: Wikipedia-based WTQ\textbf{WTQ}, financial TAT-QA\textbf{TAT-QA}, and scientific SCITAB\textbf{SCITAB}. Although instruction tuning and larger, newer LLMs deliver stronger, more robust TQA performance, data contamination and reliability issues, especially on WTQ\textbf{WTQ}, remain unresolved. Through an in-depth attention analysis, we reveal a strong correlation between perturbation-induced shifts in attention dispersion and the drops in performance, with sensitivity peaking in the model's middle layers. We highlight the need for improved interpretable methodologies to develop more reliable LLMs for table comprehension. Through an in-depth attention analysis, we reveal a strong correlation between perturbation-induced shifts in attention dispersion and performance drops, with sensitivity peaking in the model's middle layers. Based on these findings, we argue for the development of structure-aware self-attention mechanisms and domain-adaptive processing techniques to improve the transparency, generalization, and real-world reliability of LLMs on tabular data.

Keywords

Cite

@article{arxiv.2406.12719,
  title  = {Exploring the Robustness of Language Models for Tabular Question Answering via Attention Analysis},
  author = {Kushal Raj Bhandari and Sixue Xing and Soham Dan and Jianxi Gao},
  journal= {arXiv preprint arXiv:2406.12719},
  year   = {2025}
}

Comments

Accepted TMLR 2025