English
Related papers

Related papers: Testing the Testers: Human-Driven Quality Assessme…

200 papers

In this study, it was investigated whether AI evaluators assess the content validity of B1-level English reading comprehension test items in a manner similar to human evaluators. A 25-item multiple-choice test was developed, and these test…

Human-Computer Interaction · Computer Science 2025-03-21 Hatice Gurdil , Hatice Ozlem Anadol , Yesim Beril Soguksu

Healthcare conversational AI agents shouldn't be optimized only for clean benchmark accuracy in production-first regime; they must be optimized for the lived reality of patient conversations, where audio is imperfect, intent is indirect,…

Text-based online counselling scales across geographical and stigma barriers, yet faces practitioner shortages, lacks non-verbal cues and suffers inconsistent quality assurance. Whilst artificial intelligence offers promising solutions, its…

Computers and Society · Computer Science 2026-01-15 Philipp Steigerwald , Jennifer Burghardt , Eric Rudolph , Jens Albrecht

Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world…

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language…

Machine Learning · Computer Science 2026-05-29 Sy-Tuyen Ho , Minghui Liu , Huy Nghiem , Furong Huang

This research-to-practice full paper was inspired by the persistent challenge in effective communication among engineering students. Public speaking is a necessary skill for future engineers as they have to communicate technical knowledge…

Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling evaluations…

Computation and Language · Computer Science 2026-05-21 Md Tahmid Rahman Laskar , Xue-Yong Fu , Seyyed Saeed Sarfjoo , Quinten McNamara , Jonas Robertson , Shashi Bhushan TN

The potential presented by Artificial Intelligence (AI) for healthcare has long been recognised by the technical community. More recently, this potential has been recognised by policymakers, resulting in considerable public and private…

Artificial Intelligence · Computer Science 2021-04-15 Jessica Morley , Caroline Morton , Kassandra Karpathakis , Mariarosaria Taddeo , Luciano Floridi

Software testing remains critical for ensuring reliability, yet traditional approaches are slow, costly, and prone to gaps in coverage. This paper presents an AI-driven framework that automates test case generation and validation using…

Software Engineering · Computer Science 2025-08-25 Saba Naqvi , Mohammad Baqar

Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-learning perspective that overlooks speech-specific properties…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-05 Yi-Cheng Lin , Yun-Shao Tsai , Kuan-Yu Chen , Hsiao-Ying Huang , Huang-Cheng Chou , Hung-yi Lee

Context: The rapid adoption of AI-assisted code generation tools, such as large language models (LLMs), is transforming software development practices. While these tools promise significant productivity gains, concerns regarding the…

Software Engineering · Computer Science 2026-03-27 Vehid Geruslu , Zulfiyya Aliyeva , Eray Tüzün

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to the external validity…

Machine Learning · Computer Science 2026-03-03 Luke Guerdan , Justin Whitehouse , Kimberly Truong , Kenneth Holstein , Zhiwei Steven Wu

Artificial Intelligence (AI)-powered features have rapidly proliferated across mobile apps in various domains, including productivity, education, entertainment, and creativity. However, how users perceive, evaluate, and critique these AI…

Software Engineering · Computer Science 2025-06-13 Vinaik Chhetri , Krishna Upadhyay , A. B. Siddique , Umar Farooq

With the current progress of Artificial Intelligence (AI) technology and its increasingly broader applications, trust is seen as a required criterion for AI usage, acceptance, and deployment. A robust measurement instrument is essential to…

Human-Computer Interaction · Computer Science 2025-10-27 Retno Larasati

The quality of vocal delivery is one of the key indicators for evaluating teacher enthusiasm, which has been widely accepted to be connected to the overall course qualities. However, existing evaluation for vocal delivery is mainly…

Sound · Computer Science 2021-07-19 Hang Li , Yu Kang , Yang Hao , Wenbiao Ding , Zhongqin Wu , Zitao Liu

Virtual human animations have a wide range of applications in virtual and augmented reality. While automatic generation methods of animated virtual humans have been developed, assessing their quality remains challenging. Recently,…

Graphics · Computer Science 2025-11-17 Rim Rekik , Stefanie Wuhrer , Ludovic Hoyet , Katja Zibrek , Anne-Hélène Olivier

High-fidelity, AI-based simulated classroom systems enable teachers to rehearse effective teaching strategies. However, dialogue-oriented open-ended conversations such as teaching a student about scale factors can be difficult to model.…

Human-Computer Interaction · Computer Science 2021-12-07 Debajyoti Datta , Maria Phillips , James P Bywater , Jennifer Chiu , Ginger S. Watson , Laura E. Barnes , Donald E Brown

The scholarly publishing ecosystem faces a dual crisis of unmanageable submission volumes and unregulated AI, creating an urgent need for new governance models to safeguard scientific integrity. The traditional human-only peer review regime…

Artificial Intelligence · Computer Science 2025-10-03 Khalid M. Saqr

AI-based systems, including Large Language Models (LLM), impact millions by supporting diverse tasks but face issues like misinformation, bias, and misuse. AI ethics is crucial as new technologies and concerns emerge, but objective,…

Computers and Society · Computer Science 2025-05-19 José Antonio Siqueira de Cerqueira , Mamia Agbese , Rebekah Rousi , Nannan Xi , Juho Hamari , Pekka Abrahamsson

Artificial intelligence develops techniques and systems whose performance must be evaluated on a regular basis in order to certify and foster progress in the discipline. We will describe and critically assess the different ways AI systems…

Artificial Intelligence · Computer Science 2016-08-23 Jose Hernandez-Orallo