English
Related papers

Related papers: The Reasonable Effectiveness of Diverse Evaluation…

200 papers

Providing rich, constructive feedback to students is essential for supporting and enhancing their learning. Recent advancements in Generative Artificial Intelligence (AI), particularly with large language models (LLMs), present new…

Computers and Society · Computer Science 2025-07-11 Euan D Lindsay , Mike Zhang , Aditya Johri , Johannes Bjerva

Generative AI tools are increasingly entering academic peer review workflows, raising questions about fairness, accountability, and the legitimacy of evaluative judgment. While these systems promise efficiency gains amid growing reviewer…

Computers and Society · Computer Science 2026-03-24 Tatiana Chakravorti , Pranav Narayanan Venkit , Sourojit Ghosh , Sarah Rajtmajer

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model…

Computation and Language · Computer Science 2025-03-03 Matthias Orlikowski , Paul Röttger , Philipp Cimiano , Dirk Hovy

Generative AI chatbots have proven surprisingly effective at persuading people to change their beliefs and attitudes in lab settings. However, the practical implications of these findings are not yet clear. In this work, we explore the…

Human-Computer Interaction · Computer Science 2026-01-29 Jeremy Foote , Deepak Kumar , Bedadyuti Jha , Ryan Funkhouser , Loizos Bitsikokos , Hitesh Goel , Hsuen-Chi Chiu

The rise of generative AI (GenAI) has impacted many aspects of human life. As these systems become embedded in everyday practices, understanding public trust in them is also essential for responsible adoption and governance. Prior work on…

Computation and Language · Computer Science 2026-03-25 Aria Pessianzadeh , Naima Sultana , Hildegarde Van den Bulck , David Gefen , Shahin Jabbari , Rezvaneh Rezapour

Feedback in creativity support tools can help crowdworkers to improve their ideations. However, current feedback methods require human assessment from facilitators or peers. This is not scalable to large crowds. We propose Interpretable…

Human-Computer Interaction · Computer Science 2022-03-29 Yunlong Wang , Priyadarshini Venkatesh , Brian Y. Lim

We conducted controlled experimental bias audits for four versions of ChatGPT, which we asked to recommend an opening offer in salary negotiations for a new hire. We submitted 98,800 prompts to each version, systematically varying the…

Computers and Society · Computer Science 2024-10-10 R. Stuart Geiger , Flynn O'Sullivan , Elsie Wang , Jonathan Lo

Generative Artificial Intelligence (GenAI) has prompted significant discussion in education, yet large-scale empirical evidence on how students and teachers perceive and navigate this shift remains limited. We analyse 270k AI-related Reddit…

Computers and Society · Computer Science 2026-05-19 Pelin Yüce , Xiangruo Dai , Rebecca Owens , Tuğrulcan Elmas

Humans quite frequently interact with conversational agents. The rapid advancement in generative language modeling through neural networks has helped advance the creation of intelligent conversational agents. Researchers typically evaluate…

Computation and Language · Computer Science 2020-02-27 Sashank Santhanam , Alireza Karduni , Samira Shaikh

Human ratings are one of the most prevalent methods to evaluate the performance of natural language processing algorithms. Similarly, it is common to measure the quality of sentences generated by a natural language generation model using…

Computation and Language · Computer Science 2021-04-13 Jakob Nyberg , Ramesh Manuvinakurike , Maike Paetzel-Prüsmann

Generative AI systems such as ChatGPT and Claude are built upon language models that are typically evaluated for accuracy on curated benchmark datasets. Such evaluation paradigms measure predictive and reasoning capabilities of language…

Human-Computer Interaction · Computer Science 2025-03-03 Shreya Rajagopal , Jae Ho Sohn , Hari Subramonyam , Shiwali Mohan

Integrating generative AI such as Large Language Models into social robots has improved their ability to engage in natural, human-like communication. This study presents a method to examine their persuasive capabilities. We designed an…

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

Artificial Intelligence · Computer Science 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

Generative AI models continue to become more powerful. The launch of ChatGPT in November 2022 has ushered in a new era of AI. ChatGPT and other similar chatbots have a range of capabilities, from answering student homework questions to…

Computers and Society · Computer Science 2023-06-14 Yanchen Wang , Lisa Singh

Prevailing methods for assessing and comparing generative AIs incentivize responses that serve a hypothetical representative individual. Evaluating models in these terms presumes homogeneous preferences across the population and engenders…

Machine Learning · Computer Science 2023-03-06 Dilip Arumugam , Shi Dong , Benjamin Van Roy

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to…

Computation and Language · Computer Science 2025-06-13 Tommaso Giorgi , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

Generative artificial intelligence (GenAI) offers promising potential for advancing human-AI collaboration in qualitative research. However, existing works focused on conventional machine-learning and pattern-based AI systems, and little is…

In collaboration with Postpartum Support International (PSI), a non-profit organization dedicated to supporting caregivers with postpartum mood and anxiety disorders, we developed three chatbots to provide context-specific empathetic…

Computation and Language · Computer Science 2023-08-16 Xuewen Yao , Miriam Mikhelson , S. Craig Watkins , Eunsol Choi , Edison Thomaz , Kaya de Barbaro

Generative AI is rapidly transforming how organizations create value and evaluate talent. While large language models enhance baseline output quality, they simultaneously introduce ambiguity in assessing human creativity, as observable…

Human-Computer Interaction · Computer Science 2026-04-23 Yigal Rosen , Ilia Rushkin
‹ Prev 1 3 4 5 6 7 10 Next ›