English

All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing

Computation and Language 2026-04-28 v3

Abstract

Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity typing tasks where the space of labels is extremely large. In this work, we explore the limitations of the knowledge acquired by PLMs by proposing a novel heuristic to approximate the pre-training distribution of entities when the pre-training data is unknown. Then, we systematically demonstrate that entity-typing approaches that rely solely on the parametric knowledge of PLMs struggle significantly with entities at the long tail of the pre-training distribution, and that knowledge-infused approaches can account for some of these shortcomings. Our findings suggest that we need to go beyond PLMs to produce solutions that perform well for infrequent entities.

Keywords

Cite

@article{arxiv.2410.17355,
  title  = {All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing},
  author = {Advait Deshmukh and Ashwin Umadi and Dananjay Srinivas and Maria Leonor Pacheco},
  journal= {arXiv preprint arXiv:2410.17355},
  year   = {2026}
}