English

Domain Adaptive Pretraining for Multilingual Acronym Extraction

Computation and Language 2022-07-01 v1

Abstract

This paper presents our findings from participating in the multilingual acronym extraction shared task SDU@AAAI-22. The task consists of acronym extraction from documents in 6 languages within scientific and legal domains. To address multilingual acronym extraction we employed BiLSTM-CRF with multilingual XLM-RoBERTa embeddings. We pretrained the XLM-RoBERTa model on the shared task corpus to further adapt XLM-RoBERTa embeddings to the shared task domain(s). Our system (team: SMR-NLP) achieved competitive performance for acronym extraction across all the languages.

Cite

@article{arxiv.2206.15221,
  title  = {Domain Adaptive Pretraining for Multilingual Acronym Extraction},
  author = {Usama Yaseen and Stefan Langer},
  journal= {arXiv preprint arXiv:2206.15221},
  year   = {2022}
}

Comments

SDU@AAAI-22

R2 v1 2026-06-24T12:09:34.501Z