English
Related papers

Related papers: The Reasonable Effectiveness of Diverse Evaluation…

200 papers

While diversity has become a debated issue in design, very little research exists on positive use-cases for diversity beyond scholarly criticism. The current work addresses this gap through the case of a diversity-aware chatbot, exploring…

Human-Computer Interaction · Computer Science 2024-02-14 Peter Kun , Amalia De Götzen , Miriam Bidoglia , Niels Jørgen Gommesen , George Gaskell

As text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer. Prior…

Computation and Language · Computer Science 2022-12-27 Liam Dugan , Daphne Ippolito , Arun Kirubarajan , Sherry Shi , Chris Callison-Burch

In NLP annotation, it is common to have multiple annotators label the text and then obtain the ground truth labels based on the agreement of major annotators. However, annotators are individuals with different backgrounds, and minors'…

Computation and Language · Computer Science 2023-01-13 Ruyuan Wan , Jaehyung Kim , Dongyeop Kang

Language generation models' democratization benefits many domains, from answering health-related questions to enhancing education by providing AI-driven tutoring services. However, language generation models' democratization also makes it…

Computation and Language · Computer Science 2021-06-03 Paras Bhatt , Anthony Rios

Autoregressive language models, which use deep learning to produce human-like texts, have become increasingly widespread. Such models are powering popular virtual assistants in areas like smart health, finance, and autonomous driving. While…

Artificial Intelligence · Computer Science 2023-03-16 Kaiping Chen , Anqi Shao , Jirayu Burapacheep , Yixuan Li

Large Language Models (LLMs), which simulate human users, are frequently employed to evaluate chatbots in applications such as tutoring and customer service. Effective evaluation necessitates a high degree of human-like diversity within…

Computation and Language · Computer Science 2024-09-04 Xiaoyu Lin , Xinkai Yu , Ankit Aich , Salvatore Giorgi , Lyle Ungar

Human-annotated data plays a critical role in the fairness of AI systems, including those that deal with life-altering decisions or moderating human-created web/social media content. Conventionally, annotator disagreements are resolved…

Information Retrieval · Computer Science 2023-07-21 Tharindu Cyril Weerasooriya , Sarah Luger , Saloni Poddar , Ashiqur R. KhudaBukhsh , Christopher M. Homan

The launch of ChatGPT in November 2022 marked the beginning of a new era in AI, the availability of generative AI tools for everyone to use. ChatGPT and other similar chatbots boast a wide range of capabilities from answering student…

Computation and Language · Computer Science 2024-09-04 Yanchen Wang , Lisa Singh

This paper documents a collaborative research process involving peacebuilders and data scientists in Kenya and Sudan to develop AI-based text classifiers for monitoring online polarization and hatespeech. The method describes a…

This paper reports on an audit study of generative AI systems (ChatGPT, Bing Chat, and Perplexity) which investigates how these new search engines construct responses and establish authority for topics of public importance. We collected…

Information Retrieval · Computer Science 2024-05-24 Alice Li , Luanne Sinnamon

In recent months, the social impact of Artificial Intelligence (AI) has gained considerable public interest, driven by the emergence of Generative AI models, ChatGPT in particular. The rapid development of these models has sparked heated…

Artificial Intelligence · Computer Science 2024-05-10 Maria T. Baldassarre , Danilo Caivano , Berenice Fernandez Nieto , Domenico Gigante , Azzurra Ragone

Recent studies suggest that while generative AI (GenAI) can enhance individual creativity, it often reduces the diversity of collective outputs. A well-known example of this homogenization effect is by Doshi and Hauser (2024) who found that…

Human-Computer Interaction · Computer Science 2026-03-26 Yun Wan , Yoram M Kalman

The rise of generative AI (GenAI) chatbots accessible via conversational interfaces is transforming digital interactions and holds economic promise. However, these tools might deepen existing inequalities -- not only through uneven,…

A rapidly increasing amount of human conversation occurs online. But divisiveness and conflict can fester in text-based interactions on social media platforms, in messaging apps, and on other digital forums. Such toxicity increases…

Human-Computer Interaction · Computer Science 2023-10-24 Lisa P. Argyle , Ethan Busby , Joshua Gubler , Chris Bail , Thomas Howe , Christopher Rytting , David Wingate

Grounding conversations in existing passages, known as Retrieval-Augmented Generation (RAG), is an important aspect of Chat-Based Assistants powered by Large Language Models (LLMs) to ensure they are faithful and don't provide…

Human-Computer Interaction · Computer Science 2025-10-15 Sara Rosenthal , Maeda Hanafi , Yannis Katsis , Lucian Popa , Marina Danilevsky

Demographic information is often used to model annotator perspectives in subjective tasks such as hate speech detection, but its benefit is inconsistent: it improves performance in some settings and behaves as noise in others. This paper…

Computation and Language · Computer Science 2026-05-27 Weibin Cai , Reza Zafarani

The evidence on the effects of generative AI (GenAI) on critical thinking is mixed, with studies suggesting both potential harms and benefits depending on its implementation. Some argue that AI-driven provocations, such as questions asking…

Human-Computer Interaction · Computer Science 2026-03-23 Thomas Şerban von Davier , Hao-Ping Lee , Jodi Forlizzi , Sauvik Das

What are the limits of automated Twitter sentiment classification? We analyze a large set of manually labeled tweets in different languages, use them as training data, and construct automated classification models. It turns out that the…

Computation and Language · Computer Science 2021-08-31 Igor Mozetic , Miha Grcar , Jasmina Smailovic

As AI advances in text generation, human trust in AI generated content remains constrained by biases that go beyond concerns of accuracy. This study explores how bias shapes the perception of AI versus human generated content. Through three…

Computation and Language · Computer Science 2025-08-07 Tiffany Zhu , Iain Weissburg , Kexun Zhang , William Yang Wang

Summarization systems are ultimately evaluated by human annotators and raters. Usually, annotators and raters do not reflect the demographics of end users, but are recruited through student populations or crowdsourcing platforms with skewed…

Computation and Language · Computer Science 2021-10-12 Anna Jørgensen , Anders Søgaard