English

TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection

Computation and Language 2018-02-21 v1

Abstract

Detecting novelty of an entire document is an Artificial Intelligence (AI) frontier problem that has widespread NLP applications, such as extractive document summarization, tracking development of news events, predicting impact of scholarly articles, etc. Important though the problem is, we are unaware of any benchmark document level data that correctly addresses the evaluation of automatic novelty detection techniques in a classification framework. To bridge this gap, we present here a resource for benchmarking the techniques for document level novelty detection. We create the resource via event-specific crawling of news documents across several domains in a periodic manner. We release the annotated corpus with necessary statistics and show its use with a developed system for the problem in concern.

Keywords

Cite

@article{arxiv.1802.06950,
  title  = {TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection},
  author = {Tirthankar Ghosal and Amitra Salam and Swati Tiwari and Asif Ekbal and Pushpak Bhattacharyya},
  journal= {arXiv preprint arXiv:1802.06950},
  year   = {2018}
}

Comments

Accepted for publication in Language Resources and Evaluation Conference (LREC) 2018

R2 v1 2026-06-23T00:27:11.825Z