English

Financial Numeric Extreme Labelling: A Dataset and Benchmarking for XBRL Tagging

Computation and Language 2023-06-07 v1 Artificial Intelligence Computational Engineering, Finance, and Science

Abstract

The U.S. Securities and Exchange Commission (SEC) mandates all public companies to file periodic financial statements that should contain numerals annotated with a particular label from a taxonomy. In this paper, we formulate the task of automating the assignment of a label to a particular numeral span in a sentence from an extremely large label set. Towards this task, we release a dataset, Financial Numeric Extreme Labelling (FNXL), annotated with 2,794 labels. We benchmark the performance of the FNXL dataset by formulating the task as (a) a sequence labelling problem and (b) a pipeline with span extraction followed by Extreme Classification. Although the two approaches perform comparably, the pipeline solution provides a slight edge for the least frequent labels.

Cite

@article{arxiv.2306.03723,
  title  = {Financial Numeric Extreme Labelling: A Dataset and Benchmarking for XBRL Tagging},
  author = {Soumya Sharma and Subhendu Khatuya and Manjunath Hegde and Afreen Shaikh. Koustuv Dasgupta and Pawan Goyal and Niloy Ganguly},
  journal= {arXiv preprint arXiv:2306.03723},
  year   = {2023}
}

Comments

Accepted to ACL'23 Findings Paper

R2 v1 2026-06-28T10:57:52.218Z