English
Related papers

Related papers: Quantitatively ranking incorrect responses to mult…

200 papers

Non-Factoid (NF) Question Answering (QA) is challenging to evaluate due to diverse potential answers and no objective criterion. The commonly used automatic evaluation metrics like ROUGE or BERTScore cannot accurately measure semantic…

Computation and Language · Computer Science 2024-10-01 Sihui Yang , Keping Bi , Wanqing Cui , Jiafeng Guo , Xueqi Cheng

An important, yet largely unstudied, problem in student data analysis is to detect misconceptions from students' responses to open-response questions. Misconception detection enables instructors to deliver more targeted feedback on the…

Machine Learning · Statistics 2017-03-31 Joshua J. Michalenko , Andrew S. Lan , Richard G. Baraniuk

As part of large-scale assessment project at Texas Tech University, we studied the effect of problem format on students responses to quiz questions. The same problem was written in multiple formats and administered as a quiz in the large…

Physics Education · Physics 2013-12-23 Beth Thacker , Ganesh Chapagain , David Pattillo , Keith West

Modeling item parameters as a function of item characteristics has a long history but has generally focused on models for item location. Explanatory item response models for item discrimination are available but rarely used. In this study,…

Methodology · Statistics 2025-06-24 Joshua B. Gilbert , Lijin Zhang , Esther Ulitzsch , Benjamin W. Domingue

Measurement bridges theory and empirics. Without measures that appropriately capture theoretical concepts, description will fail to represent reality and true causal inference will be impossible. Yet, the social sciences traffic in complex…

Applications · Statistics 2024-05-29 Marco Morucci , Margaret Foster , Kaitlyn Webster , So Jin Lee , David Siegel

We study the problem of answering questions about images in the harder setting, where the test questions and corresponding images contain novel objects, which were not queried about in the training data. Such setting is inevitable in real…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Santhosh K. Ramakrishnan , Ambar Pal , Gaurav Sharma , Anurag Mittal

Student responses in STEM assessments are often handwritten and combine symbolic expressions, calculations, and diagrams, creating substantial variation in format and interpretation. Despite their importance for evaluating students'…

Artificial Intelligence · Computer Science 2026-04-15 Xiuxiu Tang , G. Alex Ambrose , Ying Cheng

Science is an inherently quantitative endeavor, and general education science courses are taken by a majority of college students. As such, they are a powerful venue for advancing students' skills and attitudes toward mathematics. This…

Physics Education · Physics 2015-07-16 Katherine B. Follette , Donald W. McCarthy , Erin Dokter , Sanlyn Buxner , Edward Prather

Computerized Adaptive Testing (CAT) is a widely used, efficient test mode that adapts to the examinee's proficiency level in the test domain. CAT requires pre-trained item profiles, for CAT iteratively assesses the student real-time based…

Machine Learning · Computer Science 2025-03-12 Soonwoo Kwon , Sojung Kim , Seunghyun Lee , Jin-Young Kim , Suyeong An , Kyuseok Kim

In multiple-choice exams, students select one answer from among typically four choices and can explain why they made that particular choice. Students are good at understanding natural language questions and based on their domain knowledge…

Computation and Language · Computer Science 2021-10-19 Jennifer D'Souza , Isaiah Onando Mulang' , Soeren Auer

Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration. This study…

Computation and Language · Computer Science 2026-01-07 Christopher Ormerod

Learning to Rank (LTR) from user interactions is challenging as user feedback often contains high levels of bias and noise. At the moment, two methodologies for dealing with bias prevail in the field of LTR: counterfactual methods that…

Information Retrieval · Computer Science 2019-07-16 Rolf Jagerman , Harrie Oosterhuis , Maarten de Rijke

Statistical models such as those derived from Item Response Theory (IRT) enable the assessment of students on a specific subject, which can be useful for several purposes (e.g., learning path customization, drop-out prediction). However,…

Computation and Language · Computer Science 2020-05-07 Luca Benedetto , Andrea Cappelli , Roberto Turrin , Paolo Cremonesi

International Large-scale Assessments (ILSAs), such as the Program for International Student Assessment (PISA) and the Trends in International Mathematics and Science Study (TIMSS), are cornerstone tools for global educational research and…

Methodology · Statistics 2026-05-05 Jing Ouyang , Yunxiao Chen , Chengcheng Li , Gongjun Xu

Language Models (LMs) have shown promising performance in natural language generation. However, as LMs often generate incorrect or hallucinated responses, it is crucial to correctly quantify their uncertainty in responding to given inputs.…

Computation and Language · Computer Science 2024-09-17 Xinmeng Huang , Shuo Li , Mengxin Yu , Matteo Sesia , Hamed Hassani , Insup Lee , Osbert Bastani , Edgar Dobriban

The Natural Questions (NQ) benchmark set brings new challenges to Machine Reading Comprehension: the answers are not only at different levels of granularity (long and short), but also of richer types (including no-answer, yes/no,…

Computation and Language · Computer Science 2020-09-30 Xuguang Wang , Linjun Shou , Ming Gong , Nan Duan , Daxin Jiang

Counterfactual reasoning is an important paradigm applicable in many fields, such as healthcare, economics, and education. In this work, we propose a novel method to address the issue of \textit{selection bias}. We learn two groups of…

Machine Learning · Computer Science 2019-12-20 Zichen Zhang , Qingfeng Lan , Lei Ding , Yue Wang , Negar Hassanpour , Russell Greiner

Long chain-of-thought (CoT) trajectories provide rich supervision signals for distilling reasoning from teacher to student LLMs. However, both prior work and our experiments show that trajectories from stronger teachers do not necessarily…

Research suggests "write-to-learn" tasks improve learning outcomes, yet constructed-response methods of formative assessment become unwieldy with large class sizes. This study evaluates natural language processing algorithms to assist this…

Other Statistics · Statistics 2023-01-30 Susan Lloyd , Matthew Beckman , Dennis Pearl , Rebecca Passonneau , Zhaohui Li , Zekun Wang

The Force Concept Inventory (FCI) has been widely used to assess student understanding of introductory mechanics concepts by a variety of educators and physics education researchers. One reason for this extensive use is that many of the…

Physics Education · Physics 2016-06-24 Alexandru Maries , Chandralekha Singh