English
Related papers

Related papers: The Use of Artificial Intelligence Tools in Assess…

200 papers

Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena…

Recent research has focused on literary machine translation (MT) as a new challenge in MT. However, the evaluation of literary MT remains an open problem. We contribute to this ongoing discussion by introducing LITEVAL-CORPUS, a…

Computation and Language · Computer Science 2025-02-26 Ran Zhang , Wei Zhao , Steffen Eger

Automatic verification deals with the validation by means of computers of correctness certificates. The related tools, usually called proof assistants or interactive provers, provide an interactive environment for the creation of formal…

Logic in Computer Science · Computer Science 2017-01-16 Andrea Asperti

Artificial Intelligence (AI) and large language models (LLMs) are increasingly used in social and psychological research. Among potential applications, LLMs can be used to generate, customise, or adapt measurement instruments. This study…

Human-Computer Interaction · Computer Science 2026-02-16 Mario Angelelli , Morena Oliva , Serena Arima , Enrico Ciavolino

This paper examines how individuals perceive the credibility of content originating from human authors versus content generated by large language models, like the GPT language model family that powers ChatGPT, in different user interface…

Human-Computer Interaction · Computer Science 2023-09-07 Martin Huschens , Martin Briesch , Dominik Sobania , Franz Rothlauf

Generative AI systems have rapidly advanced, with multimodal input capabilities enabling reasoning beyond text-based tasks. In education, these advancements could influence assessment design and question answering, presenting both…

Computers and Society · Computer Science 2025-07-08 Aymeric de Chillaz , Anna Sotnikova , Patrick Jermann , Antoine Bosselut

Existing LLM-as-a-Judge approaches for evaluating text generation suffer from rating inconsistencies, with low agreement and high rating variance across different evaluator models. We attribute this to subjective evaluation criteria…

Computation and Language · Computer Science 2025-11-04 Yukyung Lee , Joonghoon Kim , Jaehee Kim , Hyowon Cho , Jaewook Kang , Pilsung Kang , Najoung Kim

The large size and complex decision mechanisms of state-of-the-art text classifiers make it difficult for humans to understand their predictions, leading to a potential lack of trust by the users. These issues have led to the adoption of…

Artificial intelligence systems increasingly generate text intended to provide social and emotional support. Understanding how users perceive empathic qualities in such content is therefore critical. We examined differences in perceived…

Computers and Society · Computer Science 2026-02-20 Jonas Festor , Ivo Snels , Bennett Kleinberg

Testing Deep Learning (DL) based systems inherently requires large and representative test sets to evaluate whether DL systems generalise beyond their training datasets. Diverse Test Input Generators (TIGs) have been proposed to produce…

Software Engineering · Computer Science 2022-12-23 Vincenzo Riccio , Paolo Tonella

Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world…

Five thousand variations of the RoBERTa model, an artificially intelligent "transformer" that can understand text language, completed an English literacy exam with 29 multiple-choice questions. Data were used to calculate the psychometric…

Computation and Language · Computer Science 2023-10-19 Hotaka Maeda

Writing is a foundational literacy skill that underpins effective communication, fosters critical thinking, facilitates learning across disciplines, and enables individuals to organize and articulate complex ideas. Consequently, writing…

Computation and Language · Computer Science 2026-03-05 Jiangang Hao

Conventional methods for visual assessment of civil infrastructures have certain limitations, such as subjectivity of the collected data, long inspection time, and high cost of labor. Although some new technologies i.e. robotic techniques…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Enes Karaaslan , Ulas Bagci , F. Necati Catbas

This late-breaking work presents a large-scale analysis of explainable AI (XAI) literature to evaluate claims of human explainability. We collaborated with a professional librarian to identify 18,254 papers containing keywords related to…

Human-Computer Interaction · Computer Science 2025-03-24 Ashley Suh , Isabelle Hurley , Nora Smith , Ho Chit Siu

We examine whether a leading AI system GPT4 understands text as well as humans do, first using a well-established standardized test of discourse comprehension. On this test, GPT4 performs slightly, but not statistically significantly,…

Computation and Language · Computer Science 2025-01-22 Thomas R. Shultz , Jamie M. Wise , Ardavan Salehi Nobandegani

Language tests measure a person's ability to use a language in terms of listening, speaking, reading, or writing. Such tests play an integral role in academic, professional, and immigration domains, with entities such as educational…

Computers and Society · Computer Science 2023-07-20 Dawen Zhang , Thong Hoang , Shidong Pan , Yongquan Hu , Zhenchang Xing , Mark Staples , Xiwei Xu , Qinghua Lu , Aaron Quigley

As artificial intelligence (AI) systems approach and surpass expert human performance across a broad range of tasks, obtaining high-quality human supervision for evaluation and training becomes increasingly challenging. Our focus is on…

Machine Learning · Computer Science 2026-02-25 Ren Yin , Takashi Ishida , Masashi Sugiyama

Multimedia or spoken content presents more attractive information than plain text content, but it's more difficult to display on a screen and be selected by a user. As a result, accessing large collections of the former is much more…

Computation and Language · Computer Science 2016-08-24 Bo-Hsiang Tseng , Sheng-Syun Shen , Hung-Yi Lee , Lin-Shan Lee

Reliable evaluation of AI systems remains a fundamental challenge when ground truth labels are unavailable, particularly for systems generating natural language outputs like AI chat and agent systems. Many of these AI agents and systems…

Machine Learning · Statistics 2025-11-05 Kaihua Ding