English

A Hybrid Word-Character Approach to Abstractive Summarization

Computation and Language 2018-09-11 v2

Abstract

Automatic abstractive text summarization is an important and challenging research topic of natural language processing. Among many widely used languages, the Chinese language has a special property that a Chinese character contains rich information comparable to a word. Existing Chinese text summarization methods, either adopt totally character-based or word-based representations, fail to fully exploit the information carried by both representations. To accurately capture the essence of articles, we propose a hybrid word-character approach (HWC) which preserves the advantages of both word-based and character-based representations. We evaluate the advantage of the proposed HWC approach by applying it to two existing methods, and discover that it generates state-of-the-art performance with a margin of 24 ROUGE points on a widely used dataset LCSTS. In addition, we find an issue contained in the LCSTS dataset and offer a script to remove overlapping pairs (a summary and a short text) to create a clean dataset for the community. The proposed HWC approach also generates the best performance on the new, clean LCSTS dataset.

Keywords

Cite

@article{arxiv.1802.09968,
  title  = {A Hybrid Word-Character Approach to Abstractive Summarization},
  author = {Chieh-Teng Chang and Chi-Chia Huang and Chih-Yuan Yang and Jane Yung-Jen Hsu},
  journal= {arXiv preprint arXiv:1802.09968},
  year   = {2018}
}
R2 v1 2026-06-23T00:35:20.024Z