Related papers: Does it matter if you answer slowly?
Adaptive learning systems can produce substantial learning gains, yet many students engage for too brief or too superficial a period to benefit. A central obstacle is measuring effort. Effort during multi-step problem solving is rarely…
LLMs have demonstrated impressive zero-shot performance on NLP tasks thanks to the knowledge they acquired in their training. In multiple-choice QA tasks, the LM probabilities are used as an imperfect measure of the plausibility of each…
In this research study, we empirically investigate the effect of sampling temperature on the performance of Large Language Models (LLMs) on various problem-solving tasks. We created a multiple-choice question-and-answer (MCQA) exam by…
Grades provide students with their primary performance feedback: signals which affect academic choices. Variations in grading practice among courses impose grade penalties (and bonuses) on students who take them. These grade penalties are…
Creating equitable performance outcomes among students is a focus of many instructors and researchers. One focus of this effort is examining disparities in physics student performance across genders, which is a well-established problem.…
We describe a retrospective study of the responses to the Brief Electricity and Magnetism Assessment (BEMA) collected from a large population of 3480 students at a large public university. Two different online testing setups were employed…
Educational assessments are valuable tools for measuring student knowledge and skills, but their validity can be compromised when test takers exhibit changes in response behavior due to factors such as time pressure. To address this issue,…
The wording of natural language prompts has been shown to influence the performance of large language models (LLMs), yet the role of politeness and tone remains underexplored. In this study, we investigate how varying levels of prompt…
Automated scoring of student responses to open-ended questions, including short-answer questions, has great potential to scale to a large number of responses. Recent approaches for automated scoring rely on supervised learning, i.e.,…
This paper shows that the timing of monetary transfers to low-income families affects students' cognitive performance on high-stakes standardized tests. We combine administrative records from the world's largest conditional cash transfer…
Difficulty adjustment in practice exercises has been shown to be beneficial for learning. However, previous research has mostly investigated close-ended tasks, which do not offer the students multiple ways to reach a valid solution.…
Responsiveness in large language model (LLM) applications is widely assumed to be critical, yet the impact of latency on user behavior and perception of output quality has not been systematically explored. We report a controlled experiment…
There is growing interest in the introduction of Einsteinian concepts of space, time, light and gravity across the entire school curriculum. We have developed intervention programs and measured their effectiveness in terms of student…
When discriminating dynamic noisy sensory signals, human and primate subjects achieve higher accuracy when they take more time to decide, an effect attributed to accumulation of evidence over time to overcome neural noise. We measured the…
Despite the precision and adaptiveness of generative AI (GAI)-powered feedback provided to students, existing practice and literature might ignore how usage patterns impact student learning. This study examines the heterogeneous effects of…
Active learning strategies have been widely recognised for their effectiveness in tertiary education, yet their implementation at scale, particularly in large first-year mathematics courses, presents considerable challenges. A common method…
We discuss the development of a research-based conceptual multiple-choice survey related to magnetism. We also discuss the use of the survey to investigate gender differences in students' difficulties with concepts related to magnetism. We…
This paper suggests a generalized distribution of response times to new information $\sim t^{-b}$ for human populations in the absence of deadlines. This has important implications for psychological and social studies as well the study of…
Large language models (LLMs) often exhibit strong biases, e.g, against women or in favor of the number 7. We investigate whether LLMs would be able to output less biased answers when allowed to observe their prior answers to the same…
There has been an increasing trend of females performing better than males academically across the mathematical engineering courses. To confirm this assumption, final marks of two independent samples of students from Calculus courses across…