English
Related papers

Related papers: The Reasonable Effectiveness of Diverse Evaluation…

200 papers

Stylistic variation is critical to render the utterances generated by conversational agents natural and engaging. In this paper, we focus on sequence-to-sequence models for open-domain dialogue response generation and propose a new method…

Computation and Language · Computer Science 2018-10-02 Yujie Xing , Raquel Fernández

Most online information sources are text-based and in Western Languages like English. However, many new and first time users of the Internet are in contexts with low English proficiency and are unable to access vital information online.…

Human-Computer Interaction · Computer Science 2021-11-29 Anurag Aribandi , Divyanshu Agrawal , Dipanjan Chakraborty

Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that users must critically evaluate AI outputs rather than accept them at face value. The present…

Artificial Intelligence · Computer Science 2026-05-27 Gabriel R. Lau , Wei Yan Low , Louis Tay , Ysabel Guevarra , Dragan Gašević , Andree Hartanto

Efficiently evaluating the performance of text-to-image models is difficult as it inherently requires subjective judgment and human preference, making it hard to compare different models and quantify the state of the art. Leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Dimitrios Christodoulou , Mads Kuhlmann-Jørgensen

Generative AI (GAI) technologies are quickly reshaping the educational landscape. As adoption accelerates, understanding how students and educators perceive these tools is essential. This study presents one of the most comprehensive…

Social and Information Networks · Computer Science 2026-01-07 Paulina DeVito , Akhil Vallala , Sean Mcmahon , Yaroslav Hinda , Benjamin Thaw , Hanqi Zhuang , Hari Kalva

Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduce semantic entropy, a measure of variability across multiple…

Artificial Intelligence · Computer Science 2025-08-07 Karrtik Iyer , Manikandan Ravikiran , Prasanna Pendse , Shayan Mohanty

Traditional automated metrics for evaluating conditional natural language generation use pairwise comparisons between a single generated text and the best-matching gold-standard ground truth text. When multiple ground truths are available,…

Computation and Language · Computer Science 2022-09-30 David M Chan , Yiming Ni , David A Ross , Sudheendra Vijayanarasimhan , Austin Myers , John Canny

Over a billion users globally interact with AI systems engineered to mimic human traits. This development raises concerns that anthropomorphism, the attribution of human characteristics to AI, may foster over-reliance and misplaced trust.…

Artificial Intelligence · Computer Science 2026-02-24 Robin Schimmelpfennig , Mark Díaz , Vinodkumar Prabhakaran , Aida Davani

Open-domain chatbots are supposed to converse freely with humans without being restricted to a topic, task or domain. However, the boundaries and/or contents of open-domain conversations are not clear. To clarify the boundaries of…

Computation and Language · Computer Science 2022-11-28 A. Seza Doğruöz , Gabriel Skantze

Artificial Intelligence (AI) systems, especially generative AI technologies are becoming more relevant in our society. Tools like ChatGPT are being used by members of the disabled community e.g., Autistic people may use it to help compose…

Human-Computer Interaction · Computer Science 2023-09-27 Deepak Giri , Erin Brady

AI models have become extremely popular and accessible to the general public. However, they are continuously under the scanner due to their demonstrable biases toward various sections of the society like people of color and non-binary…

Computers and Society · Computer Science 2023-10-11 Siddharth D Jaiswal , Ankit Kumar Verma , Animesh Mukherjee

Emotions play a significant role in teamwork and collaborative activities like software development. While researchers have analyzed developer emotions in various software artifacts (e.g., issues, pull requests), few studies have focused on…

Software Engineering · Computer Science 2023-11-09 Amirali Sajadi , Kostadin Damevski , Preetha Chatterjee

The rapid advancements in large language models and generative artificial intelligence (AI) capabilities are making their broad application in the high-stakes testing context more likely. Use of generative AI in the scoring of constructed…

Computation and Language · Computer Science 2026-03-23 Jodi M. Casabianca , Daniel F. McCaffrey , Matthew S. Johnson , Naim Alper , Vladimir Zubenko

Generative AI, such as large language models, has undergone rapid development within recent years. As these models become increasingly available to the public, concerns arise about perpetuating and amplifying harmful biases in applications.…

Computation and Language · Computer Science 2024-09-04 Sara Sterlie , Nina Weng , Aasa Feragen

For machine learning datasets to accurately represent diverse opinions in a population, they must preserve variation in data labels while filtering out spam or low-quality responses. How can we balance annotator reliability and…

Computation and Language · Computer Science 2025-11-07 Eve Fleisig , Matthias Orlikowski , Philipp Cimiano , Dan Klein

This study explores linguistic differences between human and LLM-generated dialogues, using 19.5K dialogues generated by ChatGPT-3.5 as a companion to the EmpathicDialogues dataset. The research employs Linguistic Inquiry and Word Count…

Computation and Language · Computer Science 2024-04-29 Morgan Sandler , Hyesun Choung , Arun Ross , Prabu David

Evaluation of open-domain dialogue systems is highly challenging and development of better techniques is highlighted time and again as desperately needed. Despite substantial efforts to carry out reliable live evaluation of systems in…

Computation and Language · Computer Science 2022-03-14 Tianbo Ji , Yvette Graham , Gareth J. F. Jones , Chenyang Lyu , Qun Liu

As generative AI (GenAI) is increasingly applied in persona development to represent real users, understanding the implications and limitations of this technology is essential for establishing robust practices. This scoping review analyzes…

Human-Computer Interaction · Computer Science 2026-04-20 Danial Amin , Joni Salminen , Farhan Ahmed , Sonja M. H. Tervola , Sankalp Sethi , Bernard J. Jansen

Social impact evaluations are emerging as a useful tool to understand, document, and evaluate the societal impacts of generative AI. In this provocation, we begin to think carefully about the types of experts and expertise that are needed…

Human-Computer Interaction · Computer Science 2024-11-12 Zoe Kahn , Nitin Kohli

Creating a linguistic resource is often done by using a machine learning model that filters the content that goes through to a human annotator, before going into the final resource. However, budgets are often limited, and the amount of…

Computation and Language · Computer Science 2018-07-19 Filip Klubička , Giancarlo D. Salton , John D. Kelleher