English
Related papers

Related papers: Does it matter if you answer slowly?

200 papers

An important aspect of multiple hypothesis testing is controlling the significance level, or the level of Type I error. When the test statistics are not independent it can be particularly challenging to deal with this problem, without…

Statistics Theory · Mathematics 2009-03-04 Sandy Clarke , Peter Hall

The increasing availability of large language models (LLMs) has raised concerns about their potential misuse in online learning. While tools for detecting LLM-generated text exist and are widely used by researchers and educators, their…

Human-Computer Interaction · Computer Science 2025-06-23 Shambhavi Bhushan , Danielle R Thomas , Conrad Borchers , Isha Raghuvanshi , Ralph Abboud , Erin Gatz , Shivang Gupta , Kenneth Koedinger

Decisions are often made by heterogeneous groups of individuals, each with distinct initial biases and access to information of different quality. We show that in large groups of independent agents who accumulate evidence the first to…

Physics and Society · Physics 2024-01-03 Samantha Linn , Sean D. Lawley , Bhargav R. Karamched , Zachary P. Kilpatrick , Krešimir Josić

Large Language Models (LLMs) have been evaluated using diverse question types, e.g., multiple-choice, true/false, and short/long answers. This study answers an unexplored question about the impact of different question types on LLM accuracy…

Computation and Language · Computer Science 2026-04-29 Seok Hwan Song , Mohna Chakraborty , Qi Li , Wallapak Tavanapong

Language technologies have a racial bias, committing greater errors for Black users than for white users. However, little work has evaluated what effect these disparate error rates have on users themselves. The present study aims to…

Human-Computer Interaction · Computer Science 2023-02-27 Kimi Wenzel , Nitya Devireddy , Cam Davidson , Geoff Kaufman

When foraging for information, users face a tradeoff between the accuracy and value of the acquired information and the time spent collecting it, a problem which also surfaces when seeking answers to a question posed to a large community.…

Computers and Society · Computer Science 2010-08-31 Christina Aperjis , Bernardo A. Huberman , Fang Wu

This study investigates the effects of nail penetration speed on the safety outcomes of large-format automotive lithium-ion pouch cells. Through six controlled tests varying the speed of nail insertion, we observed that lower penetration…

Systems and Control · Electrical Eng. & Systems 2026-03-11 Eymen Ipek , Oliver Korak , Georg Gsellmann , Andrey Golubkov

Written responses can provide a wealth of data in understanding student reasoning on a topic. Yet they are time- and labor-intensive to score, requiring many instructors to forego them except as limited parts of summative assessments at the…

Artificial Intelligence · Computer Science 2018-05-08 Michael J Wiser , Louise S Mead , James J Smith , Robert T Pennock

In observational studies of discrimination, the most common statistical approaches consider either the rate at which decisions are made (benchmark tests) or the success rate of those decisions (outcome tests). Both tests, however, have…

Applications · Statistics 2025-03-07 Johann D. Gaebler , Sharad Goel

While research on applications and evaluations of explanation methods continues to expand, fairness of the explanation methods concerning disparities in their performance across subgroups remains an often overlooked aspect. In this paper,…

Computation and Language · Computer Science 2025-05-05 Mahdi Dhaini , Ege Erdogan , Nils Feldhus , Gjergji Kasneci

The question of how systems respond to perturbations is ubiquitous in physics. Predicting this response for large classes of systems becomes particularly challenging if many degrees of freedom are involved and linear response theory cannot…

Statistical Mechanics · Physics 2024-01-10 Lennart Dabelow , Peter Reimann

Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide how risk changes between prompt and response. We present a…

Computation and Language · Computer Science 2026-05-21 Mengya Hu , Qiong Wei , Sandeep Atluri

Reasoning models (e.g., DeepSeek-R1) generate long chains of thought to solve harder problems, but they often loop, repeating the same text at low temperatures or with greedy decoding. We study why this happens and what role temperature…

Automated short-answer scoring lags other LLM applications. We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies, modeling the traditional effect size of Quadratic Weighted Kappa (QWK) with…

Computation and Language · Computer Science 2026-03-27 Michael Hardy

In this paper, we investigate the impact of objects on gender bias in image captioning systems. Our results show that only gender-specific objects have a strong gender bias (e.g., women-lipstick). In addition, we propose a visual…

Computation and Language · Computer Science 2023-11-21 Ahmed Sabir , Lluís Padró

Response time-delay is an ubiquitous phenomenon in biological systems. Here we use a simple stochastic population model with time-delayed switching-rate conversion to quantitatively study the biological influence of the response time-delay…

Populations and Evolution · Quantitative Biology 2009-04-14 Xiao chuan Xue , Jinhua Zhao , Fei Liu , Zhong-can Ou-Yang

Tests for paired censored outcomes have been extensively studied, with some justified in the context of randomization-based inference. These tests are primarily designed to detect an overall treatment effect across the entire follow-up…

Methodology · Statistics 2025-06-10 Sangjin Lee , Kwonsang Lee

Standardized math assessments require expensive human pilot studies to establish the difficulty of test items. We investigate the predictive value of open-source large language models (LLMs) for evaluating the difficulty of multiple-choice…

Computation and Language · Computer Science 2026-04-22 Christabel Acquaye , Yi Ting Huang , Marine Carpuat , Rachel Rudinger

The use of Large Language Models (LLMs) is proliferating, yet their performance is observed to vary based on prompting styles and tones. In this study, we investigate both whether and how tonal variations in prompts lead to disparate LLM…

Artificial Intelligence · Computer Science 2026-05-29 Om Dobariya , Akhil Kumar

In recent years, numerous studies have been published dealing with the effect of individual characteristics of pedestrians on the fundamental diagram. These studies compared cumulative data on individuals in a group homogeneous in terms of…

Physics and Society · Physics 2022-03-01 Sarah Paetzke , Maik Boltes , Armin Seyfried