English

Human experts vs. machines in taxa recognition

Machine Learning 2022-02-21 v5 Machine Learning Quantitative Methods

Abstract

The step of expert taxa recognition currently slows down the response time of many bioassessments. Shifting to quicker and cheaper state-of-the-art machine learning approaches is still met with expert scepticism towards the ability and logic of machines. In our study, we investigate both the differences in accuracy and in the identification logic of taxonomic experts and machines. We propose a systematic approach utilizing deep Convolutional Neural Nets with the transfer learning paradigm and extensively evaluate it over a multi-pose taxonomic dataset with hierarchical labels specifically created for this comparison. We also study the prediction accuracy on different ranks of taxonomic hierarchy in detail. We used support vector machine classifier as a benchmark. Our results revealed that human experts using actual specimens yield the lowest classification error (CE=6.1%\overline{CE}=6.1\%). However, a much faster, automated approach using deep Convolutional Neural Nets comes close to human accuracy (CE=11.4%\overline{CE}=11.4\%) when a typical flat classification approach is used. Contrary to previous findings in the literature, we find that for machines following a typical flat classification approach commonly used in machine learning performs better than forcing machines to adopt a hierarchical, local per parent node approach used by human taxonomic experts (CE=13.8%\overline{CE}=13.8\%). Finally, we publicly share our unique dataset to serve as a public benchmark dataset in this field.

Keywords

Cite

@article{arxiv.1708.06899,
  title  = {Human experts vs. machines in taxa recognition},
  author = {Johanna Ärje and Jenni Raitoharju and Alexandros Iosifidis and Ville Tirronen and Kristian Meissner and Moncef Gabbouj and Serkan Kiranyaz and Salme Kärkkäinen},
  journal= {arXiv preprint arXiv:1708.06899},
  year   = {2022}
}

Comments

12 pages, 6 figures, 4 tables; link to the dataset fixed

R2 v1 2026-06-22T21:21:23.893Z