English

Encoding and Understanding Astrophysical Information in Large Language Model-Generated Summaries

Computation and Language 2025-11-19 v1 Instrumentation and Methods for Astrophysics

Abstract

Large Language Models have demonstrated the ability to generalize well at many levels across domains, modalities, and even shown in-context learning capabilities. This enables research questions regarding how they can be used to encode physical information that is usually only available from scientific measurements, and loosely encoded in textual descriptions. Using astrophysics as a test bed, we investigate if LLM embeddings can codify physical summary statistics that are obtained from scientific measurements through two main questions: 1) Does prompting play a role on how those quantities are codified by the LLM? and 2) What aspects of language are most important in encoding the physics represented by the measurement? We investigate this using sparse autoencoders that extract interpretable features from the text.

Keywords

Cite

@article{arxiv.2511.14685,
  title  = {Encoding and Understanding Astrophysical Information in Large Language Model-Generated Summaries},
  author = {Kiera McCormick and Rafael Martínez-Galarza},
  journal= {arXiv preprint arXiv:2511.14685},
  year   = {2025}
}

Comments

Accepted to the Machine Learning and the Physical Sciences Workshop at NeurIPS 2025, 11 pages, 4 figures

R2 v1 2026-07-01T07:43:46.250Z