English
Related papers

Related papers: Dimensions of Generative AI Evaluation Design

200 papers

Generative AI (GenAI) models have become vital across industries, yet current evaluation methods have not adapted to their widespread use. Traditional evaluations often rely on benchmarks and fixed datasets, frequently failing to reflect…

Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult. We argue that these measurement tasks are highly…

Generative AI (GenAI) has shown remarkable capabilities in generating diverse and realistic content across different formats like images, videos, and text. In Generative AI, human involvement is essential, thus HCI literature has…

Human-Computer Interaction · Computer Science 2024-01-17 Jingyu Shi , Rahul Jain , Hyungjun Doh , Ryo Suzuki , Karthik Ramani

GenAI has gained the attention of a myriad of users in almost every profession. Its advancement has had an intense impact on education, significantly disrupting the assessment design and evaluation methodologies. Despite the potential…

Computers and Society · Computer Science 2024-05-06 Rajan Kadel , Bhupesh Kumar Mishra , Samar Shailendra , Samia Abid , Maneeha Rani , Shiva Prasad Mahato

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks…

The rapid development of generative AI (GenAI) models in computer vision necessitates effective evaluation methods to ensure their quality and fairness. Existing tools primarily focus on dataset quality assurance and model explainability,…

Human-Computer Interaction · Computer Science 2024-02-07 Tica Lin , Hanspeter Pfister , Jui-Hsien Wang

The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard for valid measurement that helps place many of the…

This paper systematically derives design dimensions for the structured evaluation of explainable artificial intelligence (XAI) approaches. These dimensions enable a descriptive characterization, facilitating comparisons between different…

Human-Computer Interaction · Computer Science 2020-09-15 Fabian Sperrle , Mennatallah El-Assady , Grace Guo , Duen Horng Chau , Alex Endert , Daniel Keim

The deployment of generative AI (GenAI) models raises significant fairness concerns, addressed in this paper through novel characterization and enforcement techniques specific to GenAI. Unlike standard AI performing specific tasks, GenAI's…

Machine Learning · Computer Science 2025-08-12 Chih-Hong Cheng , Changshun Wu , Xingyu Zhao , Saddek Bensalem , Harald Ruess

Disparities in the societal harms and impacts of Generative AI (GenAI) systems highlight the critical need for effective unfairness measurement approaches. While numerous benchmarks exist, designing valid measurements requires proper…

Computers and Society · Computer Science 2025-07-08 Kimberly Le Truong , Annette Zimmermann , Hoda Heidari

Generative Artificial Intelligence (Generative AI) holds significant promise in reshaping interactive systems design, yet its potential across the four key phases of human-centered design remains underexplored. This article addresses this…

Human-Computer Interaction · Computer Science 2024-11-06 Marie Muehlhaus , Jürgen Steimle

Generative AI technologies are growing in power, utility, and use. As generative technologies are being incorporated into mainstream applications, there is a need for guidance on how to design those applications to foster productive and…

Human-Computer Interaction · Computer Science 2023-01-16 Justin D. Weisz , Michael Muller , Jessica He , Stephanie Houde

Generative AI (genAI) tools promise productivity gains, yet miscalibrated trust and usage friction still hinder adoption. Moreover, genAI can be exclusionary, failing to adequately support diverse users. One such aspect of diversity is…

Generative AI, such as large language models, has undergone rapid development within recent years. As these models become increasingly available to the public, concerns arise about perpetuating and amplifying harmful biases in applications.…

Computation and Language · Computer Science 2024-09-04 Sara Sterlie , Nina Weng , Aasa Feragen

Generative AI (GenAI) systems are inherently non-deterministic, producing varied outputs even for identical inputs. While this variability is central to their appeal, it challenges established HCI evaluation practices that typically assume…

Human-Computer Interaction · Computer Science 2026-01-30 Hyerim Park , Khanh Huynh , Malin Eiband , Jeremy Dillmann , Sven Mayer , Michael Sedlmair

Generative artificial intelligence (GenAI) holds the potential to transform the delivery, cultivation, and evaluation of human learning. This Perspective examines the integration of GenAI as a tool for human learning, addressing its…

Human-Computer Interaction · Computer Science 2024-09-06 Lixiang Yan , Samuel Greiff , Ziwen Teuber , Dragan Gašević

Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we propose a…

Generative AI applications present unique design challenges. As generative AI technologies are increasingly being incorporated into mainstream applications, there is an urgent need for guidance on how to design user experiences that foster…

Human-Computer Interaction · Computer Science 2024-01-29 Justin D. Weisz , Jessica He , Michael Muller , Gabriela Hoefer , Rachel Miles , Werner Geyer

Generative AI systems across modalities, ranging from text (including code), image, audio, and video, have broad social impacts, but there is no official standard for means of evaluating those impacts or for which impacts should be…

Generative AI enables automated, effective manipulation at scale. Despite the growing general ethical discussion around generative AI, the specific manipulation risks remain inadequately investigated. This article outlines essential…

Computers and Society · Computer Science 2025-03-10 Michael Klenk
‹ Prev 1 2 3 10 Next ›