English
Related papers

Related papers: Integrated Testlets and the Immediate Feedback Ass…

200 papers

Research on the test structure of the Force Concept Inventory (FCI) has largely been performed with exploratory methods such as factor analysis and cluster analysis. Multi-Dimensional Item Response Theory (MIRT) provides an alternative to…

Physics Education · Physics 2018-06-20 John Stewart , Cabot Zabriskie , Seth DeVore , Gay Stewart

In this paper, we introduce the Interpretable Cross-Examination Technique (ICE-T), a novel approach that leverages structured multi-prompt techniques with Large Language Models (LLMs) to improve classification performance over zero-shot and…

Computation and Language · Computer Science 2024-05-14 Goran Muric , Ben Delay , Steven Minton

As part of large-scale assessment project at Texas Tech University, we studied the effect of problem format on students responses to quiz questions. The same problem was written in multiple formats and administered as a quiz in the large…

Physics Education · Physics 2013-12-23 Beth Thacker , Ganesh Chapagain , David Pattillo , Keith West

The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data…

Computation and Language · Computer Science 2026-03-26 Tianpeng Zheng , Zhehan Jiang , Jiayi Liu , Shicong Feng

The primary goal of this study is to develop and evaluate an innovative prompting technique, AnaQuest, for generating multiple-choice questions (MCQs) using a pre-trained large language model. In AnaQuest, the choice items are…

Computation and Language · Computer Science 2025-08-08 Machi Shimmei , Masaki Uto , Yuichiroh Matsubayashi , Kentaro Inui , Aditi Mallavarapu , Noboru Matsuda

Assessment of proficiency of the learner is an essential part of Intelligent Tutoring Systems (ITS). We use Item Response Theory (IRT) in computer-aided language learning for assessment of student ability in two contexts: in test sessions,…

Artificial Intelligence · Computer Science 2024-09-25 Jue Hou , Anisia Katinskaia , Anh-Duc Vu , Roman Yangarber

Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tasks where outputs are scored continuously rather than marked…

Computation and Language · Computer Science 2026-01-21 Esma Balkır , Alice Pernthaller , Marco Basaldella , José Hernández-Orallo , Nigel Collier

Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration. This study…

Computation and Language · Computer Science 2026-01-07 Christopher Ormerod

Conceptual tests are widely used by physics instructors to assess students' conceptual understanding and compare teaching methods. It is common to look at students' changes in their answers between a pre-test and a post-test to quantify a…

Physics Education · Physics 2015-09-15 Brahim Lamine , Jean-François Parmentier

This paper explores the use of large language models (LLMs) to score and explain short-answer assessments in K-12 science. While existing methods can score more structured math and computer science assessments, they often do not provide…

Computation and Language · Computer Science 2024-05-02 Clayton Cohn , Nicole Hutchins , Tuan Le , Gautam Biswas

In computerized adaptive testing (CAT), items (questions) are selected in real time based on the already observed responses, so that the ability of the examinee can be estimated as accurately as possible. This is typically formulated as a…

Statistics Theory · Mathematics 2015-01-08 Shiyu Wang , Georgios Fellouris , Hua-Hua Chang

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross-modal integration. However, current benchmarks…

Computation and Language · Computer Science 2026-03-04 Shunki Uebayashi , Kento Masui , Kyohei Atarashi , Han Bao , Hisashi Kashima , Naoto Inoue , Mayu Otani , Koh Takeuchi

Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform. We first reveal flaws in MCQA's format, as it struggles to: 1) test generation/subjectivity;…

Computation and Language · Computer Science 2025-06-03 Nishant Balepur , Rachel Rudinger , Jordan Lee Boyd-Graber

Effective teaching relies on knowing what students know-or think they know. Revealing student thinking is challenging. Often used because of their ease of grading, even the best multiple choice (MC) tests, those using research based…

Computers and Society · Computer Science 2024-06-12 Michael Klymkowsky , Melanie M. Cooper

Computerized Adaptive Testing (CAT) is a widely used technology for evaluating learners' proficiency in online education platforms. By leveraging prior estimates of proficiency to select questions and updating the estimates iteratively…

Information Retrieval · Computer Science 2025-12-24 Mi Tian , Kun Zhang , Fei Liu , Jinglong Li , Yuxin Liao , Chenxi Bai , Zhengtao Tan , Le Wu , Richang Hong

Multiple-choice (MC) tests are an efficient method to assess English learners. It is useful for test creators to rank candidate MC questions by difficulty during exam curation. Typically, the difficulty is determined by having human test…

Computation and Language · Computer Science 2024-04-17 Vatsal Raina , Mark Gales

As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking can be trusted is essential. We evaluate LLM-as-a-judge marking across three physics assessment formats -…

Physics Education · Physics 2026-03-17 Will Yeadon , Tom Hardy , Paul Mackay , Elise Agra

Computerized adaptive testing (CAT) refers to a form of tests that are personalized to every student/test taker. CAT methods adaptively select the next most informative question/item for each student given their responses to previous…

Machine Learning · Computer Science 2021-08-18 Aritra Ghosh , Andrew Lan

Cognitive diagnosis is a fundamental and crucial task in many educational applications, e.g., computer adaptive test and cognitive assignments. Item Response Theory (IRT) is a classical cognitive diagnosis method which can provide…

Artificial Intelligence · Computer Science 2019-12-03 Song Cheng , Qi Liu

Statistical models such as those derived from Item Response Theory (IRT) enable the assessment of students on a specific subject, which can be useful for several purposes (e.g., learning path customization, drop-out prediction). However,…

Computation and Language · Computer Science 2020-05-07 Luca Benedetto , Andrea Cappelli , Roberto Turrin , Paolo Cremonesi