English

LBMT team at VLSP2022-Abmusu: Hybrid method with text correlation and generative models for Vietnamese multi-document summarization

Computation and Language 2023-04-12 v1

Abstract

Multi-document summarization is challenging because the summaries should not only describe the most important information from all documents but also provide a coherent interpretation of the documents. This paper proposes a method for multi-document summarization based on cluster similarity. In the extractive method we use hybrid model based on a modified version of the PageRank algorithm and a text correlation considerations mechanism. After generating summaries by selecting the most important sentences from each cluster, we apply BARTpho and ViT5 to construct the abstractive models. Both extractive and abstractive approaches were considered in this study. The proposed method achieves competitive results in VLSP 2022 competition.

Keywords

Cite

@article{arxiv.2304.05205,
  title  = {LBMT team at VLSP2022-Abmusu: Hybrid method with text correlation and generative models for Vietnamese multi-document summarization},
  author = {Tan-Minh Nguyen and Thai-Binh Nguyen and Hoang-Trung Nguyen and Hai-Long Nguyen and Tam Doan Thanh and Ha-Thanh Nguyen and Thi-Hai-Yen Vuong},
  journal= {arXiv preprint arXiv:2304.05205},
  year   = {2023}
}

Comments

In Proceedings of the 9th International Workshop on Vietnamese Language and Speech Processing (VLSP 2022)

R2 v1 2026-06-28T09:59:39.588Z