English

TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data

Computation and Language 2020-05-19 v1 Machine Learning

Abstract

Recent years have witnessed the burgeoning of pretrained language models (LMs) for text-based natural language (NL) understanding tasks. Such models are typically trained on free-form NL text, hence may not be suitable for tasks like semantic parsing over structured data, which require reasoning over both free-form NL questions and structured tabular data (e.g., database tables). In this paper we present TaBERT, a pretrained LM that jointly learns representations for NL sentences and (semi-)structured tables. TaBERT is trained on a large corpus of 26 million tables and their English contexts. In experiments, neural semantic parsers using TaBERT as feature representation layers achieve new best results on the challenging weakly-supervised semantic parsing benchmark WikiTableQuestions, while performing competitively on the text-to-SQL dataset Spider. Implementation of the model will be available at http://fburl.com/TaBERT .

Keywords

Cite

@article{arxiv.2005.08314,
  title  = {TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data},
  author = {Pengcheng Yin and Graham Neubig and Wen-tau Yih and Sebastian Riedel},
  journal= {arXiv preprint arXiv:2005.08314},
  year   = {2020}
}

Comments

To Appear at ACL 2020

R2 v1 2026-06-23T15:36:28.736Z