English

An Empirical Study for Vietnamese Constituency Parsing with Pre-training

Computation and Language 2020-10-21 v2

Abstract

In this work, we use a span-based approach for Vietnamese constituency parsing. Our method follows the self-attention encoder architecture and a chart decoder using a CKY-style inference algorithm. We present analyses of the experiment results of the comparison of our empirical method using pre-training models XLM-Roberta and PhoBERT on both Vietnamese datasets VietTreebank and NIIVTB1. The results show that our model with XLM-Roberta archived the significantly F1-score better than other pre-training models, VietTreebank at 81.19% and NIIVTB1 at 85.70%.

Cite

@article{arxiv.2010.09623,
  title  = {An Empirical Study for Vietnamese Constituency Parsing with Pre-training},
  author = {Tuan-Vi Tran and Xuan-Thien Pham and Duc-Vu Nguyen and Kiet Van Nguyen and Ngan Luu-Thuy Nguyen},
  journal= {arXiv preprint arXiv:2010.09623},
  year   = {2020}
}
R2 v1 2026-06-23T19:27:31.450Z