English
Related papers

Related papers: The Reasonable Effectiveness of Diverse Evaluation…

200 papers

If large language models like GPT-3 preferably produce a particular point of view, they may influence people's opinions on an unknown scale. This study investigates whether a language-model-powered writing assistant that generates some…

Human-Computer Interaction · Computer Science 2023-02-02 Maurice Jakesch , Advait Bhat , Daniel Buschek , Lior Zalmanson , Mor Naaman

Conversation agents, commonly referred to as chatbots, are increasingly deployed in many domains to allow people to have a natural interaction while trying to solve a specific problem. Given their widespread use, it is important to provide…

Social and Information Networks · Computer Science 2020-10-13 Biplav Srivastava , Francesca Rossi , Sheema Usmani , and Mariana Bernagozzi

Writing survey questions that easily and accurately convey their intent to a variety of respondents is a demanding and high-stakes task. Despite the extensive literature on best practices, the number of considerations to keep in mind is…

Methodology · Statistics 2025-09-11 Erica Ann Metheney , Lauren Yehle

As AI becomes more integral in our lives, the need for transparency and responsibility grows. While natural language explanations (NLEs) are vital for clarifying the reasoning behind AI decisions, evaluating them through human judgments is…

Computation and Language · Computer Science 2024-03-27 Fan Huang , Haewoon Kwak , Kunwoo Park , Jisun An

Label aggregation such as majority voting is commonly used to resolve annotator disagreement in dataset creation. However, this may disregard minority values and opinions. Recent studies indicate that learning from individual annotations…

Computation and Language · Computer Science 2023-10-24 Xinpeng Wang , Barbara Plank

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such as operational failures (annotators…

Artificial Intelligence · Computer Science 2026-05-08 Alex Oesterling , Donghao Ren , Yannick Assogba , Dominik Moritz , Sunnie S. Y. Kim , Leon Gatys , Fred Hohman

Recommendation systems increasingly depend on massive human-labeled datasets; however, the human annotators hired to generate these labels increasingly come from homogeneous backgrounds. This poses an issue when downstream predictive models…

Although pre-trained sequence-to-sequence models have achieved great success in dialogue response generation, chatbots still suffer from generating inconsistent responses in real-world practice, especially in multi-turn settings. We argue…

Computation and Language · Computer Science 2022-03-08 Leyang Cui , Fandong Meng , Yijin Liu , Jie Zhou , Yue Zhang

With the growing prevalence of large language models, it is increasingly common to annotate datasets for machine learning using pools of crowd raters. However, these raters often work in isolation as individual crowdworkers. In this work,…

Computers and Society · Computer Science 2024-08-05 Sonja Schmer-Galunder , Ruta Wheelock , Scott Friedman , Alyssa Chvasta , Zaria Jalan , Emily Saltz

New systems employ Machine Learning to sift through large knowledge sources, creating flexible Large Language Models. These models discern context and predict sequential information in various communication forms. Generative AI, leveraging…

Artificial Intelligence · Computer Science 2023-07-19 Ted Selker

All AI models are susceptible to learning biases in data that they are trained on. For generative dialogue models, being trained on real human conversations containing unbalanced gender and race/ethnicity references can lead to models that…

Computation and Language · Computer Science 2021-09-09 Eric Michael Smith , Adina Williams

In commonsense generation, given a set of input concepts, a model must generate a response that is not only commonsense bearing, but also capturing multiple diverse viewpoints. Numerous evaluation metrics based on form- and content-level…

Computation and Language · Computer Science 2025-06-03 Tianhui Zhang , Bei Peng , Danushka Bollegala

Aiming for a mixbiotic society that combines freedom and solidarity among people with diverse values, I focused on nonviolent communication (NVC) that enables compassionate giving in various situations of social division and conflict, and…

Artificial Intelligence · Computer Science 2023-08-08 Takeshi Kato

One challenge for evaluating current sequence- or dialogue-level chatbots, such as Empathetic Open-domain Conversation Models, is to determine whether the chatbot performs in an emotionally consistent way. The most recent work only…

Computation and Language · Computer Science 2021-12-06 Chenxiao Liu , Guanzhi Deng , Tao Ji , Difei Tang , Silai Zheng

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

People are increasingly turning to generative AI (e.g., ChatGPT, Gemini, Copilot) for emotional support and companionship. While trust is likely to play a central role in enabling these informal and unsupervised interactions, we still lack…

Human-Computer Interaction · Computer Science 2026-01-26 Riccardo Volpato , Simone Stumpf , Lisa DeBruine

How diverse are the outputs of large language models when diversity is desired? We examine the diversity of responses of various models to questions with multiple possible answers, comparing them with human responses. Our findings suggest…

Computation and Language · Computer Science 2024-11-06 Michal Shur-Ofry , Bar Horowitz-Amsalem , Adir Rahamim , Yonatan Belinkov

Data annotation remains the sine qua non of machine learning and AI. Recent empirical work on data annotation has begun to highlight the importance of rater diversity for fairness, model performance, and new lines of research have begun to…

Artificial Intelligence · Computer Science 2024-02-13 Andrew Smart , Ding Wang , Ellis Monk , Mark Díaz , Atoosa Kasirzadeh , Erin Van Liemt , Sonja Schmer-Galunder

A widespread view is that Artificial Intelligence cannot be creative. We tested this assumption by comparing human-generated ideas with those generated by six Generative Artificial Intelligence (GAI) chatbots: $alpa.\!ai$, $Copy.\!ai$,…

Artificial Intelligence · Computer Science 2023-10-18 Jennifer Haase , Paul H. P. Hanel

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only possible to gather a small amount of high-quality labels.…

Machine Learning · Computer Science 2021-10-05 Neel Nanda , Jonathan Uesato , Sven Gowal
‹ Prev 1 4 5 6 7 8 10 Next ›