English
Related papers

Related papers: SPHERE: An Evaluation Card for Human-AI Systems

200 papers

This paper explores the advancements in making large language models (LLMs) more human-like. We focus on techniques that enhance natural language understanding, conversational coherence, and emotional intelligence in AI systems. The study…

Computation and Language · Computer Science 2026-02-03 Ethem Yağız Çalık , Talha Rüzgar Akkuş

AI model documentation is fragmented across platforms and inconsistent in structure, preventing policymakers, auditors, and users from reliably assessing safety claims, data provenance, and version-level changes. We analyzed documentation…

Artificial Intelligence · Computer Science 2025-12-16 Akhmadillo Mamirov , Faiaz Azmain , Hanyu Wang

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on large-scale real-world data such as electronic health records…

Building trustworthy AI systems for mental health support is a shared priority across stakeholders from multiple disciplines. However, "trustworthy" remains loosely defined and inconsistently operationalized. AI research often focuses on…

Computation and Language · Computer Science 2026-04-23 Xin Sun , Yue Su , Yifan Mo , Qingyu Meng , Yuxuan Li , Saku Sugawara , Mengyuan Zhang , Charlotte Gerritsen , Sander L. Koole , Koen Hindriks , Jiahuan Pei

Emotion Support Conversation (ESC) is a crucial application, which aims to reduce human stress, offer emotional guidance, and ultimately enhance human mental and physical well-being. With the advancement of Large Language Models (LLMs),…

Computation and Language · Computer Science 2024-10-29 Haiquan Zhao , Lingyu Li , Shisong Chen , Shuqi Kong , Jiaan Wang , Kexin Huang , Tianle Gu , Yixu Wang , Wang Jian , Dandan Liang , Zhixu Li , Yan Teng , Yanghua Xiao , Yingchun Wang

Digital educational environments are expanding toward complex AI and human discourse, providing researchers with an abundance of data that offers deep insights into learning and instructional processes. However, traditional qualitative…

At the heart of improving conversational AI is the open problem of how to evaluate conversations. Issues with automatic metrics are well known (Liu et al., 2016, arXiv:1603.08023), with human evaluations still considered the gold standard.…

Computation and Language · Computer Science 2022-01-14 Eric Michael Smith , Orion Hsu , Rebecca Qian , Stephen Roller , Y-Lan Boureau , Jason Weston

The goal of the present paper is to develop and validate a questionnaire to assess AI literacy. In particular, the questionnaire should be deeply grounded in the existing literature on AI literacy, should be modular (i.e., including…

Artificial Intelligence · Computer Science 2023-02-21 Astrid Carolus , Martin Koch , Samantha Straka , Marc Erich Latoschik , Carolin Wienrich

The impact of using artificial intelligence (AI) to guide patient care or operational processes is an interplay of the AI model's output, the decision-making protocol based on that output, and the capacity of the stakeholders involved to…

Memory systems address the challenge of context loss in Large Language Model during prolonged interactions. However, compared to human cognition, the efficacy of these systems in processing emotion-related information remains inconclusive.…

Computation and Language · Computer Science 2026-03-02 Peng Liu , Zhen Tao , Jihao Zhao , Ding Chen , Yansong Zhang , Cuiping Li , Zhiyu Li , Hong Chen

Advances in large language models (LLMs) are profoundly reshaping the field of human-robot interaction (HRI). While prior work has highlighted the technical potential of LLMs, few studies have systematically examined their human-centered…

Robotics · Computer Science 2026-02-18 Yufeng Wang , Yuan Xu , Anastasia Nikolova , Yuxuan Wang , Jianyu Wang , Chongyang Wang , Xin Tong

Information retrieval (IR) evaluation remains challenging due to incomplete IR benchmark datasets that contain unlabeled relevant chunks. While LLMs and LLM-human hybrid strategies reduce costly human effort, they remain prone to LLM…

Computation and Language · Computer Science 2026-02-09 Minjeong Ban , Jeonghwan Choi , Hyangsuk Min , Nicole Hee-Yeon Kim , Minseok Kim , Jae-Gil Lee , Hwanjun Song

Nowadays, Artificial Intelligence (AI), particularly Machine Learning (ML) and Large Language Models (LLMs), is widely applied across various contexts. However, the corresponding models often operate as black boxes, leading them to…

Software Engineering · Computer Science 2025-12-17 Chaima Boufaied , Thanh Nguyen , Ronnie de Souza Santos

Although large language models (LLMs) have demonstrated impressive potential on simple tasks, their breadth of scope, lack of transparency, and insufficient controllability can make them less effective when assisting humans on more complex…

Human-Computer Interaction · Computer Science 2022-03-21 Tongshuang Wu , Michael Terry , Carrie J. Cai

In this paper, we develop the position that current frameworks for evaluating emotional intelligence (EI) in artificial intelligence (AI) systems need refinement because they do not adequately or comprehensively measure the various aspects…

Artificial Intelligence · Computer Science 2025-12-30 Max Parks , Kheli Atluru , Meera Vinod , Mike Kuniavsky , Jud Brewer , Sean White , Sarah Adler , Wendy Ju

Explainable Artificial Intelligence (XAI) plays a crucial role in enhancing the transparency and accountability of AI models, particularly in natural language processing (NLP) tasks. However, popular XAI methods such as LIME and SHAP have…

Artificial Intelligence · Computer Science 2024-08-19 Haoran Zheng , Utku Pamuksuz

AI governance frameworks increasingly rely on audits, yet the results of their underlying evaluations require interpretation and context to be meaningfully informative. Even technically rigorous evaluations can offer little useful insight…

Computers and Society · Computer Science 2025-08-18 Leon Staufer , Mick Yang , Anka Reuel , Stephen Casper

AI compliance is becoming increasingly critical as AI systems grow more powerful and pervasive. Yet the rapid expansion of AI policies creates substantial burdens for resource-constrained practitioners lacking policy expertise. Existing…

Human-Computer Interaction · Computer Science 2026-03-26 Yu Yang , Ig-Jae Kim , Dongwook Yoon

Empathetic Conversational Systems (ECS) are built to respond empathetically to the user's emotions and sentiments, regardless of the application domain. Current ECS studies evaluation approaches are restricted to offline evaluation…

Computation and Language · Computer Science 2024-07-29 Aravind Sesagiri Raamkumar , Siyuan Brandon Loh