English

Evaluating the Performance of Large Language Models for Spanish Language in Undergraduate Admissions Exams

Computation and Language 2023-12-29 v1 Artificial Intelligence

Abstract

This study evaluates the performance of large language models, specifically GPT-3.5 and BARD (supported by Gemini Pro model), in undergraduate admissions exams proposed by the National Polytechnic Institute in Mexico. The exams cover Engineering/Mathematical and Physical Sciences, Biological and Medical Sciences, and Social and Administrative Sciences. Both models demonstrated proficiency, exceeding the minimum acceptance scores for respective academic programs to up to 75% for some academic programs. GPT-3.5 outperformed BARD in Mathematics and Physics, while BARD performed better in History and questions related to factual information. Overall, GPT-3.5 marginally surpassed BARD with scores of 60.94% and 60.42%, respectively.

Keywords

Cite

@article{arxiv.2312.16845,
  title  = {Evaluating the Performance of Large Language Models for Spanish Language in Undergraduate Admissions Exams},
  author = {Sabino Miranda and Obdulia Pichardo-Lagunas and Bella Martínez-Seis and Pierre Baldi},
  journal= {arXiv preprint arXiv:2312.16845},
  year   = {2023}
}

Comments

11 pages, 1 figure. Submitted to a journal