English

Boosting Low-Resource Biomedical QA via Entity-Aware Masking Strategies

Computation and Language 2021-02-17 v1 Information Retrieval Machine Learning

Abstract

Biomedical question-answering (QA) has gained increased attention for its capability to provide users with high-quality information from a vast scientific literature. Although an increasing number of biomedical QA datasets has been recently made available, those resources are still rather limited and expensive to produce. Transfer learning via pre-trained language models (LMs) has been shown as a promising approach to leverage existing general-purpose knowledge. However, finetuning these large models can be costly and time consuming, often yielding limited benefits when adapting to specific themes of specialised domains, such as the COVID-19 literature. To bootstrap further their domain adaptation, we propose a simple yet unexplored approach, which we call biomedical entity-aware masking (BEM). We encourage masked language models to learn entity-centric knowledge based on the pivotal entities characterizing the domain at hand, and employ those entities to drive the LM fine-tuning. The resulting strategy is a downstream process applicable to a wide variety of masked LMs, not requiring additional memory or components in the neural architectures. Experimental results show performance on par with state-of-the-art models on several biomedical QA datasets.

Keywords

Cite

@article{arxiv.2102.08366,
  title  = {Boosting Low-Resource Biomedical QA via Entity-Aware Masking Strategies},
  author = {Gabriele Pergola and Elena Kochkina and Lin Gui and Maria Liakata and Yulan He},
  journal= {arXiv preprint arXiv:2102.08366},
  year   = {2021}
}

Comments

EACL 2021 - Short Paper - European Chapter of the Association for Computational Linguistics

R2 v1 2026-06-23T23:13:25.864Z