Related papers: Does it matter if you answer slowly?
A paper evaluating the effects of lessons intended to encourage high school students to continue physics studies made some important errors. One was to underestimate the width of confidence intervals by failing to use standard cluster…
While zero-shot instructional prompts like "Let's think step-by-step" have revolutionized Large Language Model performance, a fundamental question remains unanswered: which specific words drive their remarkable effectiveness? We introduce…
When assessing student work, graders will often find that some students will leave one or more problems blank on assessments. Since there is no work shown, the grader has no means to evaluate the student's understanding of a particular…
Uncertainty estimation is critical for deploying reasoning language models, yet remains poorly understood under extended chain-of-thought reasoning. We study parallel sampling as a fully black-box approach using verbalized confidence and…
The persistent underrepresentation of women and gender minorities within the physical sciences remains a significant issue. This study investigates gender dynamics in introductory algebra-based physics laboratories, focusing on…
At large institutions of higher education, students frequently have a choice whether to attend the introductory physics sequence asynchronously online, on-site in a traditional lecture-setting, or in a reformed studio setting. In this…
This study offers new insights into students' interest in computer science (CS) education by disentangling the distinct effects of age and gender across a diverse adolescent sample. Grounded in the person-object theory of interest (POI), we…
Large language models cannot estimate how long their own tasks take. We investigate this limitation through four experiments across 68 tasks and four model families. Pre-task estimates overshoot actual duration by 4--7$\times$ ($p <…
We examined how model size, temperature, and prompt style affect Large Language Models' (LLMs) alignment within itself, between models, and with human in assessing clinical reasoning skills. Model size emerged as a key factor in LLM-human…
We examine a measure of individual student gain by preservice elementary teachers, related to Richard Hakes use of mean gain in the study of reform classes in undergraduate physics. The gain statistic assesses the amount individual students…
Closed-loop or feedback control ratchets use information about the state of the system to operate with the aim of maximizing the performance of the system. In this paper we investigate the effects of a time delay in the feedback for a…
Automated scoring of student work at scale requires balancing accuracy against cost and latency. In "cascade" systems, small language models (LMs) handle easier scoring tasks while escalating harder ones to larger LMs -- but the challenge…
Test-time scaling, which is also often referred to as slow-thinking, has been demonstrated to enhance multi-step reasoning in large language models (LLMs). However, despite its widespread utilization, the mechanisms underlying slow-thinking…
Undergraduate graders are frequently important contributors to the teaching team in post-secondary education settings. This study set out to investigate agreement for a team of undergraduate graders as they acquired training and experience…
The power-law TST reaction rate coefficient for an elementary bimolecular reaction is studied when the reaction takes place in a nonequilibrium system with power-law distributions. We derive a generalized TST rate coefficient, which not…
Reasoning models have attracted increasing attention for their ability to tackle complex tasks, embodying the System II (slow thinking) paradigm in contrast to System I (fast, intuitive responses). Yet a key question remains: Does slower…
Using results of a Monte Carlo simulation of the Sherrington-Kirkpatrick model, we try to characterize the slow disorder samples, namely we analyze visually the correlation between the relaxation time for a given disorder sample $J$ with…
As part of large-scale assessment project at Texas Tech University, we studied the effect of problem format on students responses to quiz questions. The same problem was written in multiple formats and administered as a quiz in the large…
Objectives: This study aims to investigate the readability and understandability of bitwise operators in programming, with the main hypothesis that there will be a difference in the performance metrics (response time and error rate) between…
Adaptive experiments can increase the chance that current students obtain better outcomes from a field experiment of an instructional intervention. In such experiments, the probability of assigning students to conditions changes while more…