Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
Abstract
AI assistants are being increasingly used by students enrolled in higher education institutions. While these tools provide opportunities for improved teaching and education, they also pose significant challenges for assessment and learning outcomes. We conceptualize these challenges through the lens of vulnerability, the potential for university assessments and learning outcomes to be impacted by student use of generative AI. We investigate the potential scale of this vulnerability by measuring the degree to which AI assistants can complete assessment questions in standard university-level STEM courses. Specifically, we compile a novel dataset of textual assessment questions from 50 courses at EPFL and evaluate whether two AI assistants, GPT-3.5 and GPT-4 can adequately answer these questions. We use eight prompting strategies to produce responses and find that GPT-4 answers an average of 65.8% of questions correctly, and can even produce the correct answer across at least one prompting strategy for 85.1% of questions. When grouping courses in our dataset by degree program, these systems already pass non-project assessments of large numbers of core courses in various degree programs, posing risks to higher education accreditation that will be amplified as these models improve. Our results call for revising program-level assessment design in higher education in light of advances in generative AI.
Cite
@article{arxiv.2408.11841,
title = {Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants},
author = {Beatriz Borges and Negar Foroutan and Deniz Bayazit and Anna Sotnikova and Syrielle Montariol and Tanya Nazaretzky and Mohammadreza Banaei and Alireza Sakhaeirad and Philippe Servant and Seyed Parsa Neshaei and Jibril Frej and Angelika Romanou and Gail Weiss and Sepideh Mamooler and Zeming Chen and Simin Fan and Silin Gao and Mete Ismayilzada and Debjit Paul and Alexandre Schöpfer and Andrej Janchevski and Anja Tiede and Clarence Linden and Emanuele Troiani and Francesco Salvi and Freya Behrens and Giacomo Orsi and Giovanni Piccioli and Hadrien Sevel and Louis Coulon and Manuela Pineros-Rodriguez and Marin Bonnassies and Pierre Hellich and Puck van Gerwen and Sankalp Gambhir and Solal Pirelli and Thomas Blanchard and Timothée Callens and Toni Abi Aoun and Yannick Calvino Alonso and Yuri Cho and Alberto Chiappa and Antonio Sclocchi and Étienne Bruno and Florian Hofhammer and Gabriel Pescia and Geovani Rizk and Leello Dadi and Lucas Stoffl and Manoel Horta Ribeiro and Matthieu Bovel and Yueyang Pan and Aleksandra Radenovic and Alexandre Alahi and Alexander Mathis and Anne-Florence Bitbol and Boi Faltings and Cécile Hébert and Devis Tuia and François Maréchal and George Candea and Giuseppe Carleo and Jean-Cédric Chappelier and Nicolas Flammarion and Jean-Marie Fürbringer and Jean-Philippe Pellet and Karl Aberer and Lenka Zdeborová and Marcel Salathé and Martin Jaggi and Martin Rajman and Mathias Payer and Matthieu Wyart and Michael Gastpar and Michele Ceriotti and Ola Svensson and Olivier Lévêque and Paolo Ienne and Rachid Guerraoui and Robert West and Sanidhya Kashyap and Valerio Piazza and Viesturs Simanis and Viktor Kuncak and Volkan Cevher and Philippe Schwaller and Sacha Friedli and Patrick Jermann and Tanja Käser and Antoine Bosselut},
journal= {arXiv preprint arXiv:2408.11841},
year = {2024}
}
Comments
20 pages, 8 figures