English
Related papers

Related papers: The Use of Artificial Intelligence Tools in Assess…

200 papers

We analysed a dataset of scientific manuscripts that were submitted to various conferences in artificial intelligence. We performed a combination of semantic, lexical and psycholinguistic analyses of the full text of the manuscripts and…

Digital Libraries · Computer Science 2020-03-05 Philippe Vincent-Lamarre , Vincent Larivière

Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing…

Machine translation evaluation is a very important activity in machine translation development. Automatic evaluation metrics proposed in literature are inadequate as they require one or more human reference translations to compare them with…

Computation and Language · Computer Science 2013-11-18 Nisheeth Joshi , Iti Mathur , Hemant Darbari , Ajai Kumar

Automatic evaluation metrics capable of replacing human judgments are critical to allowing fast development of new methods. Thus, numerous research efforts have focused on crafting such metrics. In this work, we take a step back and analyze…

Computation and Language · Computer Science 2022-10-10 Pierre Colombo , Maxime Peyrard , Nathan Noiry , Robert West , Pablo Piantanida

This book critically analyses the value of citation data, altmetrics, and artificial intelligence to support the research evaluation of articles, scholars, departments, universities, countries, and funders. It introduces and discusses…

Digital Libraries · Computer Science 2025-04-14 Mike Thelwall

As large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, statistical bias in benchmark data and probing studies have recently called into question their true…

Computation and Language · Computer Science 2021-09-13 Shane Storks , Joyce Chai

The manual assessment and grading of student writing is a time-consuming yet critical task for teachers. Recent developments in generative AI, such as large language models, offer potential solutions to facilitate essay-scoring tasks for…

Computation and Language · Computer Science 2024-11-26 Kathrin Seßler , Maurice Fürstenberg , Babette Bühler , Enkelejda Kasneci

This study seeks to enhance academic integrity by providing tools to detect AI-generated content in student work using advanced technologies. The findings promote transparency and accountability, helping educators maintain ethical standards…

Computation and Language · Computer Science 2025-01-07 Ayat A. Najjar , Huthaifa I. Ashqar , Omar A. Darwish , Eman Hammad

Reading and evaluating product reviews is central to how most people decide what to buy and consume online. However, the recent emergence of Large Language Models and Generative Artificial Intelligence now means writing fraudulent or fake…

This paper presents a theoretical framework for addressing the challenges posed by generative artificial intelligence (AI) in higher education assessment through a machine-versus-machine approach. Large language models like GPT-4, Claude,…

Computers and Society · Computer Science 2025-06-04 Mohammad Saleh Torkestani , Taha Mansouri

Reward learning algorithms utilize human feedback to infer a reward function, which is then used to train an AI system. This human feedback is often a preference comparison, in which the human teacher compares several samples of AI behavior…

Machine Learning · Computer Science 2023-03-03 Peter Barnett , Rachel Freedman , Justin Svegliato , Stuart Russell

We conducted a systematic literature review on automated grading and feedback tools for programming education. We analysed 121 research papers from 2017 to 2021 inclusive and categorised them based on skills assessed, approach, language…

Software Engineering · Computer Science 2023-12-11 Marcus Messer , Neil C. C. Brown , Michael Kölling , Miaojing Shi

We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First, we implement a Deep Research pipeline that processes the…

Artificial Intelligence · Computer Science 2026-05-29 Gaurav Sahu , Laurent Charlin , Christopher Pal

Classic evaluation methods of believable agents are time-consuming because they involve many human to judge agents. They are well suited to validate work on new believable behaviours models. However, during the implementation, numerous…

Artificial Intelligence · Computer Science 2010-09-03 Fabien Tencé , Cédric Buche

Effective teaching relies on knowing what students know-or think they know. Revealing student thinking is challenging. Often used because of their ease of grading, even the best multiple choice (MC) tests, those using research based…

Computers and Society · Computer Science 2024-06-12 Michael Klymkowsky , Melanie M. Cooper

In this paper we analyze features to classify human- and AI-generated text for English, French, German and Spanish and compare them across languages. We investigate two scenarios: (1) The detection of text generated by AI from scratch, and…

Computation and Language · Computer Science 2024-01-31 Kristina Schaaff , Tim Schlippe , Lorenz Mindner

This study investigates whether individuals can learn to accurately discriminate between human-written and AI-produced texts when provided with immediate feedback, and if they can use this feedback to recalibrate their self-perceived…

Computation and Language · Computer Science 2025-10-17 Jiří Milička , Anna Marklová , Ondřej Drobil , Eva Pospíšilová

Explainability is widely regarded as essential for trustworthy artificial intelligence systems. However, the metrics commonly used to evaluate counterfactual explanations are algorithmic evaluation metrics that are rarely validated against…

Artificial Intelligence · Computer Science 2026-03-17 Felix Liedeker , Basil Ell , Philipp Cimiano , Christoph Düsing

High-stakes prediction tasks (e.g., patient diagnosis) are often handled by trained human experts. A common source of concern about automation in these settings is that experts may exercise intuition that is difficult to model and/or have…

Machine Learning · Statistics 2024-11-26 Rohan Alur , Loren Laine , Darrick K. Li , Manish Raghavan , Devavrat Shah , Dennis Shung

The quality of machine translation has increased remarkably over the past years, to the degree that it was found to be indistinguishable from professional human translation in a number of empirical investigations. We reassess Hassan et…

Computation and Language · Computer Science 2020-04-06 Samuel Läubli , Sheila Castilho , Graham Neubig , Rico Sennrich , Qinlan Shen , Antonio Toral
‹ Prev 1 3 4 5 6 7 10 Next ›