English
Related papers

Related papers: A baseline for multiple-choice testing in the univ…

200 papers

Large Language Models (LLMs) have displayed massive improvements in reasoning and decision-making skills and can hold natural conversations with users. Recently, many tool-use benchmark datasets have been proposed. However, existing…

Psychometric functions typically characterize binary sensory decisions along a single stimulus dimension. However, real-life sensory tasks vary along a greater variety of dimensions (e.g. color, contrast and luminance for visual stimuli).…

Neurons and Cognition · Quantitative Biology 2023-02-03 Stephen Keeley , Benjamin Letham , Chase Tymms , Craig Sanders , Michael Shvartsman

Large Language Models (LLMs) are transforming writing, reading, teaching, and knowledge retrieval in many academic fields. However, concerns regarding their misuse and erroneous outputs have led to varying degrees of trust in LLMs within…

Computers and Society · Computer Science 2025-02-10 Minseok Jung , Aurora Zhang , May Fung , Junho Lee , Paul Pu Liang

Large Language Models (LLMs) are increasingly embedded in academic writing practices. Although numerous studies have explored how researchers employ these tools for scientific writing, their concrete implementation, limitations, and design…

Human-Computer Interaction · Computer Science 2025-12-15 Brenda Nogueira , Werner Geyer , Andrew Anderson , Toby Jia-Jun Li , Dongwhi Kim , Nuno Moniz , Nitesh V. Chawla

Predicting the difficulty of multiple-choice questions (MCQs) is important for effective assessment, yet current methods typically assume a unimodal student ability distribution, overlooking the heterogeneous nature of student…

Computers and Society · Computer Science 2026-05-19 Dhriti Krishnan , Jaromir Savelka

The objective of physics laboratory training is to develop, in students, a variety of important cognitive and psycho-motor abilities related to experimental physics. These include conceptual understanding, procedural understanding,…

Physics Education · Physics 2013-11-26 Rajesh B. Khaparde

Considerable interest has recently been focused on studying multiple phenotypes simultaneously in both epidemiological and genomic studies, either to capture the multidimensionality of complex disorders or to understand shared etiology of…

Methodology · Statistics 2015-11-26 Denis Agniel , Katherine P. Liao , Tianxi Cai

The progression from novice to disciplinary expert is a longstanding area of inquiry in educational research. Studies investigating such progressions have often resorted to participants' self-assessments or other qualitative indicators as a…

Physics Education · Physics 2025-08-08 Julien-Pooya Weihs , Adrien Weihs , Vegard Gjerde , Helge Drange

The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving graduate-level and computational mathematics relatively…

Computation and Language · Computer Science 2026-03-05 Bianca Raimondi , Francesco Pivi , Davide Evangelista , Maurizio Gabbrielli

As learning machines increase their influence on decisions concerning human lives, analyzing their fairness properties becomes a subject of central importance. Yet, our best tools for measuring the fairness of learning systems are rigid…

Machine Learning · Statistics 2022-07-21 David Lopez-Paz , Diane Bouchacourt , Levent Sagun , Nicolas Usunier

A strong sense of classroom community is associated with many positive learning outcomes and is a critical contributor to undergraduate students' persistence in STEM, particularly for women and students of color. This chapter describes a…

Other Statistics · Statistics 2025-04-09 Shira Viel , Maria Tackett , Sarwari Das , Joseph Choo

Content-focused research-based assessment instruments typically use items (i.e., questions) as the unit of assessment for scoring, reporting, and validation. Couplet scoring employs an alternative unit of assessment called a couplet, which…

Physics Education · Physics 2024-12-11 Michael Vignal , Gayle Geschwind , Marcos D. Caballero , H. J. Lewandowski

We report our experiences implementing standards-based grading at scale in an Algorithms course, which serves as the terminal required CS Theory course in our department's undergraduate curriculum. The course had 200-400 students, taught by…

Computers and Society · Computer Science 2022-04-27 Lijun Chen , Joshua A. Grochow , Ryan Layer , Michael Levet

While the Question Generation (QG) task has been increasingly adopted in educational assessments, its evaluation remains limited by approaches that lack a clear connection to the educational values of test items. In this work, we introduce…

Computation and Language · Computer Science 2025-06-23 Bang Nguyen , Tingting Du , Mengxia Yu , Lawrence Angrave , Meng Jiang

Blended mathematical sensemaking in science (blended MSS) involves deep conceptual understanding of quantitative relationships describing phenomena in science and has been studies in various disciplines. However, no unified characterization…

Physics Education · Physics 2023-05-25 Leonora Kaldaras , Carl Wieman

We discuss observational studies that test many causal hypotheses, either hypotheses about many outcomes or many treatments. To be credible an observational study that tests many causal hypotheses must demonstrate that its conclusions are…

Methodology · Statistics 2017-03-08 Qingyuan Zhao , Dylan S. Small , Paul R. Rosenbaum

Many statisticians regularly teach large lecture courses on statistics, probability, or mathematics for students from other fields such as business and economics, social sciences and psychology, etc. The corresponding exams often use a…

Applications · Statistics 2025-10-06 Achim Zeileis

This short contribution reports the development of a test for assessing middle school students' physics proficiency via multiple-choice single-select items in German language. The test assesses students' content and procedural knowledge…

Physics Education · Physics 2023-03-17 Markus Sebastian Feser , Dietmar Höttecke

Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, public health, and online trust and safety. Standard approaches…

Artificial Intelligence · Computer Science 2026-04-24 Fridolin Linder , Thomas Leeper , Daniel Haimovich , Niek Tax , Lorenzo Perini , Milan Vojnovic

The aim of the work presented in this paper is to develop and evaluate an integrated system that provides automated lecture style evaluation, allowing teachers to get instant feedback related to the goodness of their lecturing style. The…

Computers and Society · Computer Science 2023-12-29 Eleni Dimitriadou , Andreas Lanitis
‹ Prev 1 3 4 5 6 7 10 Next ›