English

Experimenting with Large Language Models and vector embeddings in NASA SciX

Computation and Language 2023-12-25 v1 Instrumentation and Methods for Astrophysics Artificial Intelligence

Abstract

Open-source Large Language Models enable projects such as NASA SciX (i.e., NASA ADS) to think out of the box and try alternative approaches for information retrieval and data augmentation, while respecting data copyright and users' privacy. However, when large language models are directly prompted with questions without any context, they are prone to hallucination. At NASA SciX we have developed an experiment where we created semantic vectors for our large collection of abstracts and full-text content, and we designed a prompt system to ask questions using contextual chunks from our system. Based on a non-systematic human evaluation, the experiment shows a lower degree of hallucination and better responses when using Retrieval Augmented Generation. Further exploration is required to design new features and data augmentation processes at NASA SciX that leverages this technology while respecting the high level of trust and quality that the project holds.

Keywords

Cite

@article{arxiv.2312.14211,
  title  = {Experimenting with Large Language Models and vector embeddings in NASA SciX},
  author = {Sergi Blanco-Cuaresma and Ioana Ciucă and Alberto Accomazzi and Michael J. Kurtz and Edwin A. Henneken and Kelly E. Lockhart and Felix Grezes and Thomas Allen and Golnaz Shapurian and Carolyn S. Grant and Donna M. Thompson and Timothy W. Hostetler and Matthew R. Templeton and Shinyi Chen and Jennifer Koch and Taylor Jacovich and Daniel Chivvis and Fernanda de Macedo Alves and Jean-Claude Paquin and Jennifer Bartlett and Mugdha Polimera and Stephanie Jarmak},
  journal= {arXiv preprint arXiv:2312.14211},
  year   = {2023}
}

Comments

To appear in the proceedings of the 33th annual international Astronomical Data Analysis Software & Systems (ADASS XXXIII)

R2 v1 2026-06-28T13:59:11.345Z