English

Adaptive and Multi-Source Entity Matching for Name Standardization of Astronomical Observation Facilities

Computation and Language 2025-10-08 v1 Instrumentation and Methods for Astrophysics

Abstract

This ongoing work focuses on the development of a methodology for generating a multi-source mapping of astronomical observation facilities. To compare two entities, we compute scores with adaptable criteria and Natural Language Processing (NLP) techniques (Bag-of-Words approaches, sequential approaches, and surface approaches) to map entities extracted from eight semantic artifacts, including Wikidata and astronomy-oriented resources. We utilize every property available, such as labels, definitions, descriptions, external identifiers, and more domain-specific properties, such as the observation wavebands, spacecraft launch dates, funding agencies, etc. Finally, we use a Large Language Model (LLM) to accept or reject a mapping suggestion and provide a justification, ensuring the plausibility and FAIRness of the validated synonym pairs. The resulting mapping is composed of multi-source synonym sets providing only one standardized label per entity. Those mappings will be used to feed our Name Resolver API and will be integrated into the International Virtual Observatory Alliance (IVOA) Vocabularies and the OntoPortal-Astro platform.

Keywords

Cite

@article{arxiv.2510.05744,
  title  = {Adaptive and Multi-Source Entity Matching for Name Standardization of Astronomical Observation Facilities},
  author = {Liza Fretel and Baptiste Cecconi and Laura Debisschop},
  journal= {arXiv preprint arXiv:2510.05744},
  year   = {2025}
}

Comments

Accepted in Ontology Matching 2025 conference proceedings

R2 v1 2026-07-01T06:20:56.763Z