English

HFL at SemEval-2022 Task 8: A Linguistics-inspired Regression Model with Data Augmentation for Multilingual News Similarity

Computation and Language 2022-04-12 v1

Abstract

This paper describes our system designed for SemEval-2022 Task 8: Multilingual News Article Similarity. We proposed a linguistics-inspired model trained with a few task-specific strategies. The main techniques of our system are: 1) data augmentation, 2) multi-label loss, 3) adapted R-Drop, 4) samples reconstruction with the head-tail combination. We also present a brief analysis of some negative methods like two-tower architecture. Our system ranked 1st on the leaderboard while achieving a Pearson's Correlation Coefficient of 0.818 on the official evaluation set.

Keywords

Cite

@article{arxiv.2204.04844,
  title  = {HFL at SemEval-2022 Task 8: A Linguistics-inspired Regression Model with Data Augmentation for Multilingual News Similarity},
  author = {Zihang Xu and Ziqing Yang and Yiming Cui and Zhigang Chen},
  journal= {arXiv preprint arXiv:2204.04844},
  year   = {2022}
}

Comments

6 pages; SemEval-2022 Task 8