English

Word-level Human Interpretable Scoring Mechanism for Novel Text Detection Using Tsetlin Machines

Computation and Language 2021-05-12 v1 Artificial Intelligence Machine Learning

Abstract

Recent research in novelty detection focuses mainly on document-level classification, employing deep neural networks (DNN). However, the black-box nature of DNNs makes it difficult to extract an exact explanation of why a document is considered novel. In addition, dealing with novelty at the word-level is crucial to provide a more fine-grained analysis than what is available at the document level. In this work, we propose a Tsetlin machine (TM)-based architecture for scoring individual words according to their contribution to novelty. Our approach encodes a description of the novel documents using the linguistic patterns captured by TM clauses. We then adopt this description to measure how much a word contributes to making documents novel. Our experimental results demonstrate how our approach breaks down novelty into interpretable phrases, successfully measuring novelty.

Keywords

Cite

@article{arxiv.2105.04708,
  title  = {Word-level Human Interpretable Scoring Mechanism for Novel Text Detection Using Tsetlin Machines},
  author = {Bimal Bhattarai and Ole-Christoffer Granmo and Lei Jiao},
  journal= {arXiv preprint arXiv:2105.04708},
  year   = {2021}
}

Comments

18 pages, 11 figures, 8 tables

R2 v1 2026-06-24T01:58:03.808Z