English

Importance of Synthesizing High-quality Data for Text-to-SQL Parsing

Computation and Language 2022-12-20 v1

Abstract

Recently, there has been increasing interest in synthesizing data to improve downstream text-to-SQL tasks. In this paper, we first examined the existing synthesized datasets and discovered that state-of-the-art text-to-SQL algorithms did not further improve on popular benchmarks when trained with augmented synthetic data. We observed two shortcomings: illogical synthetic SQL queries from independent column sampling and arbitrary table joins. To address these issues, we propose a novel synthesis framework that incorporates key relationships from schema, imposes strong typing, and conducts schema-distance-weighted column sampling. We also adopt an intermediate representation (IR) for the SQL-to-text task to further improve the quality of the generated natural language questions. When existing powerful semantic parsers are pre-finetuned on our high-quality synthesized data, our experiments show that these models have significant accuracy boosts on popular benchmarks, including new state-of-the-art performance on Spider.

Keywords

Cite

@article{arxiv.2212.08785,
  title  = {Importance of Synthesizing High-quality Data for Text-to-SQL Parsing},
  author = {Yiyun Zhao and Jiarong Jiang and Yiqun Hu and Wuwei Lan and Henry Zhu and Anuj Chauhan and Alexander Li and Lin Pan and Jun Wang and Chung-Wei Hang and Sheng Zhang and Marvin Dong and Joe Lilien and Patrick Ng and Zhiguo Wang and Vittorio Castelli and Bing Xiang},
  journal= {arXiv preprint arXiv:2212.08785},
  year   = {2022}
}