English
Related papers

Related papers: Evaluation Framework for AI Systems in "the Wild"

200 papers

Through a systematization of generative AI (GenAI) stakeholder goals and expectations, this work seeks to uncover what value different stakeholders see in their contributions to the GenAI supply line. This valuation enables us to understand…

Artificial Intelligence · Computer Science 2024-08-02 Amruta Mahuli , Asia Biega

Context: Software engineering (SE) researchers increasingly study Generative AI (GenAI) while also incorporating it into their own research practices. Despite rapid adoption, there is limited empirical evidence on how GenAI is used in SE…

Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate unbiased, privacy-preserving data while maintaining…

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This…

Cryptography and Security · Computer Science 2024-10-21 Aviral Srivastava , Sourav Panda

This study investigates the optimization of Generative AI (GenAI) systems through human feedback, focusing on how varying feedback mechanisms influence the quality of GenAI outputs. We devised a Human-AI training loop where 32 students,…

Human-Computer Interaction · Computer Science 2024-04-25 Jacob Sherson , Florent Vinchon

Generative Artificial Intelligence (GenAI) is rapidly reshaping software development, with growing emphasis on accelerating productivity and optimizing performance. However, excessive focus on such dimensions risks overlooking the critical…

Software Engineering · Computer Science 2026-05-22 Mariam Guizani , Maduka Subasinghage , Sherlock A. Licorish , Sofia Ouhbi

Ensuring transparency and trust in artificial intelligence (AI) models is essential as they are increasingly deployed in safety-critical and high-stakes domains. Explainable AI (XAI) has emerged as a promising approach to address this…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Reem Hammoud , Abdul Karim Gizzini , Ali J. Ghandour

Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-AI teams are prepared to collaborate safely and effectively.…

Human-Computer Interaction · Computer Science 2026-03-20 Min Hun Lee

Generative AI has made remarkable strides to revolutionize fields such as image and video generation. These advancements are driven by innovative algorithms, architecture, and data. However, the rapid proliferation of generative models has…

Artificial Intelligence · Computer Science 2024-11-12 Dongfu Jiang , Max Ku , Tianle Li , Yuansheng Ni , Shizhuo Sun , Rongqi Fan , Wenhu Chen

As the deployment of artificial intelligence (AI) is changing many fields and industries, there are concerns about AI systems making decisions and recommendations without adequately considering various ethical aspects, such as…

Computers and Society · Computer Science 2023-10-02 Conrad Sanderson , Qinghua Lu , David Douglas , Xiwei Xu , Liming Zhu , Jon Whittle

Context. GenAI tools are being increasingly adopted by practitioners in SE, promising support for several SE activities. Despite increasing adoption, we still lack empirical evidence on how GenAI is used in practice, the benefits it…

Software Engineering · Computer Science 2026-04-02 Görkem Giray , Onur Demirörs , Marcos Kalinowski , Daniel Mendez

The widespread use of generative AI systems is coupled with significant ethical and social challenges. As a result, policymakers, academic researchers, and social advocacy groups have all called for such systems to be audited. However,…

Computers and Society · Computer Science 2024-07-09 Jakob Mokander , Justin Curl , Mihir Kshirsagar

The field of deep generative modeling has grown rapidly in the last few years. With the availability of massive amounts of training data coupled with advances in scalable unsupervised learning paradigms, recent large-scale generative models…

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning capabilities, are supported with model performance on narrow…

The use of artificial intelligence (AI) in working environments with individuals, known as Human-AI Collaboration (HAIC), has become essential in a variety of domains, boosting decision-making, efficiency, and innovation. Despite HAIC's…

Human-Computer Interaction · Computer Science 2025-03-10 George Fragiadakis , Christos Diou , George Kousiouris , Mara Nikolaidou

Generative AI tools are increasingly entering academic peer review workflows, raising questions about fairness, accountability, and the legitimacy of evaluative judgment. While these systems promise efficiency gains amid growing reviewer…

Computers and Society · Computer Science 2026-03-24 Tatiana Chakravorti , Pranav Narayanan Venkit , Sourojit Ghosh , Sarah Rajtmajer

The potential presented by Artificial Intelligence (AI) for healthcare has long been recognised by the technical community. More recently, this potential has been recognised by policymakers, resulting in considerable public and private…

Artificial Intelligence · Computer Science 2021-04-15 Jessica Morley , Caroline Morton , Kassandra Karpathakis , Mariarosaria Taddeo , Luciano Floridi

This paper investigates current software testing systems and explores how artificial intelligence, specifically Generative AI, can be integrated to enhance these systems. It begins by examining different types of AI systems and focuses on…

Software Engineering · Computer Science 2026-03-03 Tanish Singla , Qusay H. Mahmoud

In this position paper, we observe that empirical evaluation in Generative AI is at a crisis point since traditional ML evaluation and benchmarking strategies are insufficient to meet the needs of evaluating modern GenAI models and systems.…