Computation and Language · Computer Science
Introducing a framework to assess newly created questions with Natural Language Processing
Luca Benedetto, Andrea Cappelli, Roberto Turrin, Paolo Cremonesi
2020-05-07
Machine Learning · Statistics
SPRITE: A Response Model For Multiple Choice Testing
Ryan Ning, Andrew E. Waters, Christoph Studer, Richard G. Baraniuk
2015-01-14
Machine Learning · Computer Science
AutoIRT: Calibrating Item Response Theory Models with Automated Machine Learning
James Sharpnack, Phoebe Mulcaire, Klinton Bicknell, Geoff LaFlair +1
2024-09-16
Computation and Language · Computer Science
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
Xinyuan Li, Murong Xu, Wenbiao Tao, Hanlun Zhu +3
2026-02-03
Computation and Language · Computer Science
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
Sydney Peters, Nan Zhang, Hong Jiao, Ming Li +2
2025-09-30
Artificial Intelligence · Computer Science
Reasoning and Sampling-Augmented MCQ Difficulty Prediction via LLMs
Wanyong Feng, Peter Tran, Stephen Sireci, Andrew Lan
2025-03-12
Computation and Language · Computer Science
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng +3
2026-02-03
Machine Learning · Statistics
$\beta^3$-IRT: A New Item Response Model and its Applications
Yu Chen, Telmo Silva Filho, Ricardo B. C. Prudêncio, Tom Diethe +1
2019-06-04
Computation and Language · Computer Science
Difficulty-Controllable Cloze Question Distractor Generation
Seokhoon Kang, Yejin Jeon, Seonjeong Hwang, Gary Geunbae Lee
2026-05-20
Computation and Language · Computer Science
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
Christabel Acquaye, Yi Ting Huang, Marine Carpuat, Rachel Rudinger
2026-04-22
Computation and Language · Computer Science
Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking
Tianpeng Zheng, Zhehan Jiang, Jiayi Liu, Shicong Feng
2026-03-26
Computation and Language · Computer Science
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
Shunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao +4
2026-03-04
Computation and Language · Computer Science
Exploring the Potential of Large Language Models for Estimating the Reading Comprehension Question Difficulty
Yoshee Jain, John Hollander, Amber He, Sunny Tang +2
2025-02-26
Computation and Language · Computer Science
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
Zhimeng Luo, Lixin Wu, Adam Frisch, Daqing He
2026-04-07
Computation and Language · Computer Science
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen +1
2025-11-27
Computation and Language · Computer Science
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
Radhika Kapoor, Sang T. Truong, Nick Haber, Maria Araceli Ruiz-Primo +1
2026-04-01
Computation and Language · Computer Science
Reliable and Efficient Amortized Model-based Evaluation
Sang Truong, Yuheng Tu, Percy Liang, Bo Li +1
2025-03-18