English

Evaluating Byte and Wordpiece Level Models for Massively Multilingual Semantic Parsing

Computation and Language 2022-12-15 v1

Abstract

Token free approaches have been successfully applied to a series of word and span level tasks. In this work, we compare a byte-level (ByT5) and a wordpiece based (mT5) sequence to sequence model on the 51 languages of the MASSIVE multilingual semantic parsing dataset. We examine multiple experimental settings: (i) zero-shot, (ii) full gold data and (iii) zero-shot with synthetic data. By leveraging a state-of-the-art label projection method for machine translated examples, we are able to reduce the gap in exact match accuracy to only 5 points with respect to a model trained on gold data from all the languages. We additionally provide insights on the cross-lingual transfer of ByT5 and show how the model compares with respect to mT5 across all parameter sizes.

Keywords

Cite

@article{arxiv.2212.07223,
  title  = {Evaluating Byte and Wordpiece Level Models for Massively Multilingual Semantic Parsing},
  author = {Massimo Nicosia and Francesco Piccinno},
  journal= {arXiv preprint arXiv:2212.07223},
  year   = {2022}
}

Comments

Massively Multilingual NLU 2022 Workshop Paper @ EMNLP 2022 - Winning approach of the MMNLU-22 Zero-Shot Challenge

R2 v1 2026-06-28T07:34:26.632Z