English
Related papers

Related papers: Evaluation Framework for AI Systems in "the Wild"

200 papers

Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted GenAI access enables task outsourcing that undermines the validity of traditional assessments; blanket prohibitions are difficult…

Computers and Society · Computer Science 2026-05-26 Yizhu Gao , Zhongzhou Chen , Min Li , Xiaoming Zhai

Generative AI (GenAI) presents societal and ethical challenges related to equity, academic integrity, bias, and data provenance. In this paper, we outline the goals, methodology and deliverables of their collaborative research, considering…

Generative Artificial Intelligence (generative AI) poses both opportunities and risks for the integrity of research. Universities must guide researchers in using generative AI responsibly, and in navigating a complex regulatory landscape…

Computers and Society · Computer Science 2025-05-26 Shannon Smith , Melissa Tate , Keri Freeman , Anne Walsh , Brian Ballsun-Stanton , Mark Hooper , Murray Lane

As AI systems appear to exhibit ever-increasing capability and generality, assessing their true potential and safety becomes paramount. This paper contends that the prevalent evaluation methods for these systems are fundamentally…

Artificial Intelligence · Computer Science 2024-07-15 John Burden

Recent advances in Generative AI (GenAI) are transforming multiple aspects of society, including education and foreign language learning. In the context of English as a Foreign Language (EFL), significant research has been conducted to…

Computers and Society · Computer Science 2025-01-03 Jasper Roe , Mike Perkins , Leon Furze

Growing awareness of the environmental impact of digital technologies has led to several isolated initiatives to promote sustainable practices. However, despite these efforts, the environmental footprint of generative AI, particularly in…

Computers and Society · Computer Science 2024-02-06 Thomas Le Goff

As organizations grapple with the rapid adoption of Generative AI (GenAI), this study synthesizes the state of knowledge through a systematic literature review of secondary studies and research agendas. Analyzing 28 papers published since…

Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark…

Artificial Intelligence · Computer Science 2026-05-28 Aakash Pant , Kavya Shah , Apoorv Agnihotri , Sneha Nikam , Prasaanth Balraj , Nakul Jain

Generative AI systems across modalities, ranging from text (including code), image, audio, and video, have broad social impacts, but there is no official standard for means of evaluating those impacts or for which impacts should be…

Generative AI systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be measured rather than fundamental to the system's operation. Drawing on hermeneutic theory…

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based…

Artificial Intelligence · Computer Science 2024-01-01 Xiting Wang , Liming Jiang , Jose Hernandez-Orallo , David Stillwell , Luning Sun , Fang Luo , Xing Xie

Generative artificial intelligence (GenAI) holds the potential to transform the delivery, cultivation, and evaluation of human learning. This Perspective examines the integration of GenAI as a tool for human learning, addressing its…

Human-Computer Interaction · Computer Science 2024-09-06 Lixiang Yan , Samuel Greiff , Ziwen Teuber , Dragan Gašević

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this…

Societal stereotypes are at the center of a myriad of responsible AI interventions targeted at reducing the generation and propagation of potentially harmful outcomes. While these efforts are much needed, they tend to be fragmented and…

Computers and Society · Computer Science 2025-10-02 Aida Davani , Sunipa Dev , Héctor Pérez-Urbina , Vinodkumar Prabhakaran

With widespread adoption of AI models for important decision making, ensuring reliability of such models remains an important challenge. In this paper, we present an end-to-end generic framework for testing AI Models which performs…

Machine Learning · Computer Science 2021-02-12 Aniya Aggarwal , Samiulla Shaikh , Sandeep Hans , Swastik Haldar , Rema Ananthanarayanan , Diptikalyan Saha

This study examines the impact of Generative Artificial Intelligence (GenAI) on academic research, focusing on its application to qualitative and quantitative data analysis. As GenAI tools evolve rapidly, they offer new possibilities for…

Human-Computer Interaction · Computer Science 2024-08-14 Mike Perkins , Jasper Roe

Large artificial intelligence (AI) models have garnered significant attention for their remarkable, often "superhuman", performance on standardized benchmarks. However, when these models are deployed in high-stakes verticals such as…

Artificial Intelligence · Computer Science 2025-09-26 Gaurav Verma , Jiawei Zhou , Mohit Chandra , Srijan Kumar , Munmun De Choudhury

Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an…

Artificial Intelligence · Computer Science 2025-05-27 Maria Eriksson , Erasmo Purificato , Arman Noroozian , Joao Vinagre , Guillaume Chaslot , Emilia Gomez , David Fernandez-Llorca

The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance). However, divergent practices and terminologies…

Software Engineering · Computer Science 2024-05-17 Boming Xia , Qinghua Lu , Liming Zhu , Zhenchang Xing

Commonly, AI or machine learning (ML) models are evaluated on benchmark datasets. This practice supports innovative methodological research, but benchmark performance can be poorly correlated with performance in real-world applications -- a…

Machine Learning · Computer Science 2024-06-18 Olivier Binette , Jerome P. Reiter