English
Related papers

Related papers: SPHERE: An Evaluation Card for Human-AI Systems

200 papers

As artificial intelligence (AI) systems are increasingly deployed, principles for ethical AI are also proliferating. Certification offers a method to both incentivize adoption of these principles and substantiate that they have been…

Computers and Society · Computer Science 2021-05-24 Peter Cihon , Moritz J. Kleinaltenkamp , Jonas Schuett , Seth D. Baum

There has been significant recent interest in developing AI agents capable of effectively interacting and teaming with humans. While each of these works try to tackle a problem quite central to the problem of human-AI interaction, they tend…

Artificial Intelligence · Computer Science 2022-02-22 Zahra Zahedi , Sarath Sreedharan , Subbarao Kambhampati

In 2025, Large Language Model (LLM) services have launched a new feature -- AI video chat -- allowing users to interact with AI agents via real-time video communication (RTC), just like chatting with real people. Despite its significance,…

Networking and Internet Architecture · Computer Science 2025-10-02 Jiayang Xu , Xiangjie Huang , Zijie Li , Zili Meng

Communicating the risks and benefits of AI is important for regulation and public understanding. Yet current methods such as technical reports often exclude people without technical expertise. Drawing on HCI research, we developed an Impact…

Human-Computer Interaction · Computer Science 2025-08-27 Edyta Bogucka , Marios Constantinides , Sanja Šćepanović , Daniele Quercia

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human judgments can suffer…

Computation and Language · Computer Science 2019-09-24 Sashank Santhanam , Samira Shaikh

Systematic reviews are time-consuming endeavors. Historically speaking, knowledgeable humans have had to screen and extract data from studies before it can be analyzed. However, large language models (LLMs) hold promise to greatly…

Human-Computer Interaction · Computer Science 2025-01-22 Noah L. Schroeder , Chris Davis Jaldi , Shan Zhang

Customer Relationship Management (CRM) systems are vital for modern enterprises, providing a foundation for managing customer interactions and data. Integrating AI agents into CRM systems can automate routine processes and enhance…

Computation and Language · Computer Science 2025-02-18 Kung-Hsiang Huang , Akshara Prabhakar , Sidharth Dhawan , Yixin Mao , Huan Wang , Silvio Savarese , Caiming Xiong , Philippe Laban , Chien-Sheng Wu

AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable. Tools that can automate aspects of modeling practice must complement human expertise,…

Artificial Intelligence · Computer Science 2026-05-29 Sara Metcalf , William Schoenberg

As human-AI cooperation becomes increasingly prevalent, reliable instruments for assessing the subjective quality of cooperative human-AI interaction are needed. We introduce two theoretically grounded scales: the Perceived Cooperativity…

Human-Computer Interaction · Computer Science 2026-04-28 Christiane Attig , Christiane Wiebel-Herboth , Patricia Wollstadt , Tim Schrills , Mourad Zoubir , Thomas Franke

In high-stakes disaster scenarios, timely and informed decision-making is critical yet often challenged by uncertainty, dynamic environments, and limited resources. This paper presents a systematic review of Human-AI collaboration patterns…

Artificial Intelligence · Computer Science 2025-09-16 Emmanuel Adjei Domfeh , Christopher L. Dancy

As AI systems demonstrate increasingly strong predictive performance, their adoption has grown in numerous domains. However, in high-stakes domains such as criminal justice and healthcare, full automation is often not desirable due to…

Artificial Intelligence · Computer Science 2021-12-23 Vivian Lai , Chacha Chen , Q. Vera Liao , Alison Smith-Renner , Chenhao Tan

Current financial large language models (FinLLMs) struggle with two critical limitations: the absence of objective evaluation metrics to assess the quality of stock analysis reports and a lack of depth in stock analysis, which impedes their…

Artificial Intelligence · Computer Science 2025-07-10 Shijie Han , Jingshu Zhang , Yiqing Shen , Kaiyuan Yan , Hongguang Li

The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in unfair performance evaluation, often requiring manual intervention - a costly and…

Artificial Intelligence · Computer Science 2026-05-27 Jakub Slapek , Mir Seyedebrahimi , Jianhua Yang

The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for empirically evaluating anthropomorphic LLM behaviours in…

The final search query for the Systematic Literature Review (SLR) was conducted on 15th July 2022. Initially, we extracted 1707 journal and conference articles from the Scopus and Web of Science databases. Inclusion and exclusion criteria…

Artificial Intelligence · Computer Science 2022-11-29 AKM Bahalul Haque , A. K. M. Najmul Islam , Patrick Mikalef

The advancement of large language models (LLMs) has outpaced traditional evaluation methodologies. This progress presents novel challenges, such as measuring human-like psychological constructs, moving beyond static and task-specific…

Computation and Language · Computer Science 2026-03-12 Haoran Ye , Jing Jin , Yuhang Xie , Xin Zhang , Guojie Song

As the capabilities of artificial intelligence (AI) continue to expand rapidly, Human-AI (HAI) Collaboration, combining human intellect and AI systems, has become pivotal for advancing problem-solving and decision-making processes. The…

This action research study focuses on the integration of "AI assistants" in two Agile software development meetings: the Daily Scrum and a feature refinement, a planning meeting that is part of an in-house Scaled Agile framework. We discuss…

Software Engineering · Computer Science 2024-04-24 Beatriz Cabrero-Daniel , Tomas Herda , Victoria Pichler , Martin Eder

AI evaluations have become critical tools for assessing large language model capabilities and safety. This paper presents practical insights from eight months of maintaining $inspect\_evals$, an open-source repository of 70+…

Computation and Language · Computer Science 2025-07-10 Alexandra Abbas , Celia Waggoner , Justin Olive

Human feedback on conversations with language language models (LLMs) is central to how these systems learn about the world, improve their capabilities, and are steered toward desirable and safe behaviors. However, this feedback is mostly…