中文
相关论文

相关论文: An Item Response Theory-based R Module for Algorit…

200 篇论文

Semantics based knowledge representations such as ontologies are found to be very useful in automatically generating meaningful factual questions. Determining the difficulty level of these system generated questions is helpful to…

人工智能 · 计算机科学 2017-09-05 Vinu E. , P Sreenivasa Kumar

Comprehensive evaluations of language models (LM) during both development and deployment phases are necessary because these models possess numerous capabilities (e.g., mathematical reasoning, legal support, or medical diagnostic) as well as…

计算与语言 · 计算机科学 2025-03-18 Sang Truong , Yuheng Tu , Percy Liang , Bo Li , Sanmi Koyejo

We propose a class of Item Response Theory models for items with ordinal polytomous responses, which extends an existing class of multidimensional models for dichotomously-scored items measuring more than one latent trait. In the proposed…

统计方法学 · 统计学 2012-01-24 Silvia Bacci , Francesco Bartolucci , Michela Gnaldi

Item Response Theory (IRT) models have received growing interest in health science for analyzing latent constructs such as depression, anxiety, quality of life, or cognitive functioning from the information provided by each individual's…

Accuracy-based evaluation of Large Language Models (LLMs) measures benchmark-specific performance rather than underlying medical competency: it treats all questions as equally informative, conflates model ability with item characteristics,…

计算与语言 · 计算机科学 2026-04-07 Zhimeng Luo , Lixin Wu , Adam Frisch , Daqing He

The growing dependence on eTextbooks and Massive Open Online Courses (MOOCs) has led to an increase in the amount of students' learning data. By carefully analyzing this data, educators can identify difficult exercises, and evaluate the…

数据结构与算法 · 计算机科学 2022-11-28 Ahmed Abd Elrahman , Ahmed I. Taloba , Mohammed F. Farghally , Taysir Hassan A Soliman

This paper aims to present an online placement test. It is based on the Item Response Theory to provide relevant estimates of learner competences. The proposed test is the entry point of our e-Learning system. It gathers the learner…

计算机与社会 · 计算机科学 2014-11-20 Farid Merrouch , Meriem Hnida , Mohammed Khalidi Idrissi , Samir Bennani

The rapid release of both language models and benchmarks makes it increasingly costly to evaluate every model on every dataset. In practice, models are often evaluated on different samples, making scores difficult to compare across studies.…

Incorporating Item Response Theory (IRT) into NLP tasks can provide valuable information about model performance and behavior. Traditionally, IRT models are learned using human response pattern (RP) data, presenting a significant bottleneck…

计算与语言 · 计算机科学 2019-09-02 John P. Lalor , Hao Wu , Hong Yu

Item response theory (IRT) models for categorical response data are widely used in the analysis of educational data, computerized adaptive testing, and psychological surveys. However, most IRT models rely on both the assumption that…

机器学习 · 统计学 2015-01-14 Ryan Ning , Andrew E. Waters , Christoph Studer , Richard G. Baraniuk

We propose a new method to measure the task-specific accuracy of Retrieval-Augmented Large Language Models (RAG). Evaluation is performed by scoring the RAG on an automatically-generated synthetic exam composed of multiple choice questions…

计算与语言 · 计算机科学 2024-05-24 Gauthier Guinet , Behrooz Omidvar-Tehrani , Anoop Deoras , Laurent Callot

Educational assessments are valuable tools for measuring student knowledge and skills, but their validity can be compromised when test takers exhibit changes in response behavior due to factors such as time pressure. To address this issue,…

统计方法学 · 统计学 2025-05-06 Gabriel Wallin , Yunxiao Chen , Yi-Hsuan Lee , Xiaoou Li

One of the largest drivers of social inequality is unequal access to personal tutoring, with wealthier individuals able to afford it, while the majority cannot. Affordable, effective AI tutors offer a scalable solution. We focus on adaptive…

人工智能 · 计算机科学 2025-10-01 Tom Quilter , Anastasia Ilick , Karen Poon , Richard Turner

Large language models (LLMs) have demonstrated exceptional performance across a wide range of natural language tasks. However, selecting the optimal LLM to respond to a user query often necessitates a delicate balance between performance…

人工智能 · 计算机科学 2025-06-24 Wei Song , Zhenya Huang , Cheng Cheng , Weibo Gao , Bihan Xu , GuanHao Zhao , Fei Wang , Runze Wu

Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM outputs and the heterogeneity of benchmark items. Item…

机器学习 · 统计学 2026-05-11 Xinhao Qu , Qiang Heng , Hao Zeng , Xiaoqian Liu

The purpose of this study is to introduce software technologies and models and artificial intelligence algorithms to improve the weaknesses of CBT (Cognitive Behavior Therapy) method in psychotherapy. The presentation method for this…

软件工程 · 计算机科学 2021-08-17 Omid Shokrollahi

Analyses of heterogeneous treatment effects (HTE) are common in applied causal inference research. However, when outcomes are latent variables assessed via psychometric instruments such as educational tests, standard methods ignore the…

计量经济学 · 经济学 2025-06-27 Joshua B. Gilbert , Zachary Himmelsbach , James Soland , Mridul Joshi , Benjamin W. Domingue

Evaluating large language models (LLMs) on comprehensive benchmarks is a cornerstone of their development, yet it's often computationally and financially prohibitive. While Item Response Theory (IRT) offers a promising path toward…

人工智能 · 计算机科学 2025-10-07 Lele Liao , Qile Zhang , Ruofan Wu , Guanhua Fang

Generative AI is transforming the educational landscape, raising significant concerns about cheating. Despite the widespread use of multiple-choice questions in assessments, the detection of AI cheating in MCQ-based tests has been almost…

人工智能 · 计算机科学 2024-12-13 Alona Strugatski , Giora Alexandron

Item Response Theory (IRT) is widely applied in the human sciences to model persons' responses on a set of items measuring one or more latent constructs. While several R packages have been developed that implement IRT models, they tend to…

统计计算 · 统计学 2020-02-04 Paul-Christian Bürkner