English

Classification of cancer pathology reports: a large-scale comparative study

Machine Learning 2021-01-12 v1 Computation and Language Image and Video Processing Machine Learning

Abstract

We report about the application of state-of-the-art deep learning techniques to the automatic and interpretable assignment of ICD-O3 topography and morphology codes to free-text cancer reports. We present results on a large dataset (more than 80 000 labeled and 1 500 000 unlabeled anonymized reports written in Italian and collected from hospitals in Tuscany over more than a decade) and with a large number of classes (134 morphological classes and 61 topographical classes). We compare alternative architectures in terms of prediction accuracy and interpretability and show that our best model achieves a multiclass accuracy of 90.3% on topography site assignment and 84.8% on morphology type assignment. We found that in this context hierarchical models are not better than flat models and that an element-wise maximum aggregator is slightly better than attentive models on site classification. Moreover, the maximum aggregator offers a way to interpret the classification process.

Keywords

Cite

@article{arxiv.2006.16370,
  title  = {Classification of cancer pathology reports: a large-scale comparative study},
  author = {Stefano Martina and Leonardo Ventura and Paolo Frasconi},
  journal= {arXiv preprint arXiv:2006.16370},
  year   = {2021}
}

Comments

10 pages, 6 figures, 3 tables, accepted for publication in IEEE Journal of Biomedical and Health Informatics (J-BHI)

R2 v1 2026-06-23T16:42:58.615Z