English
Related papers

Related papers: Can LLMs Estimate Student Struggles? Human-AI Diff…

200 papers

It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate the confidence of their predictions. However, empirical evidence suggests that…

Machine Learning · Computer Science 2026-05-14 Nina Corvelo Benz , Eleni Straitouri , Manuel Gomez-Rodriguez

The effect of large language models (LLMs) in education is debated: Previous research shows that LLMs can help as well as hurt learning. In two pre-registered and incentivized laboratory experiments, we find no effect of LLMs on overall…

Computers and Society · Computer Science 2025-03-11 Matthias Lehmann , Philipp B. Cornelius , Fabian J. Sting

While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated…

Computation and Language · Computer Science 2026-03-02 Ali Khoramfar , Ali Ramezani , Mohammad Mahdi Mohajeri , Mohammad Javad Dousti , Majid Nili Ahmadabadi , Heshaam Faili

Machine learning (ML) and artificial intelligence (AI) tools increasingly permeate every possible social, political, and economic sphere; sorting, taxonomizing and predicting complex human behaviour and social phenomena. However, from…

Computers and Society · Computer Science 2022-06-10 Abeba Birhane

Automated short-answer scoring lags other LLM applications. We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies, modeling the traditional effect size of Quadratic Weighted Kappa (QWK) with…

Computation and Language · Computer Science 2026-03-27 Michael Hardy

As large language models (LLMs) are increasingly used to model and augment collective decision-making, it is critical to examine their alignment with human social reasoning. We present an empirical framework for assessing collective…

Artificial Intelligence · Computer Science 2025-10-03 Crystal Qian , Aaron Parisi , Clémentine Bouleau , Vivian Tsai , Maël Lebreton , Lucas Dixon

With the recent rise of widely successful deep learning models, there is emerging interest among professionals in various math and science communities to see and evaluate the state-of-the-art models' abilities to collaborate on finding or…

Computation and Language · Computer Science 2023-10-18 Sophia Gu

In many practical applications of AI, an AI model is used as a decision aid for human users. The AI provides advice that a human (sometimes) incorporates into their decision-making process. The AI advice is often presented with some measure…

Artificial Intelligence · Computer Science 2022-10-31 Kailas Vodrahalli , Tobias Gerstenberg , James Zou

Deployed artificial intelligence (AI) often impacts humans, and there is no one-size-fits-all metric to evaluate these tools. Human-centered evaluation of AI-based systems combines quantitative and qualitative analysis and human input. It…

Human-Computer Interaction · Computer Science 2023-03-14 Teresa Datta , John P. Dickerson

Large language models cannot estimate how long their own tasks take. We investigate this limitation through four experiments across 68 tasks and four model families. Pre-task estimates overshoot actual duration by 4--7$\times$ ($p <…

Computation and Language · Computer Science 2026-04-02 Aniketh Garikaparthi

Modern neural networks (NNs) often achieve high predictive accuracy but are poorly calibrated, producing overconfident predictions even when wrong. This miscalibration poses serious challenges in applications where reliable uncertainty…

Machine Learning · Computer Science 2025-09-12 Pedro Mendes , Paolo Romano , David Garlan

Large language models (LLMs) have the potential to boost human productivity by speeding up task completion -- provided users know when to offload cognitive work to them. But we do not know if users are well-calibrated in estimating these…

Computers and Society · Computer Science 2026-05-25 Sunny Yu , Myra Cheng , Ahmad Jabbar , Ilia Sucholutsky , Katherine M. Collins , Dan Jurafsky , Robert D. Hawkins

Despite the rapid expansion of Large Language Models (LLMs) in healthcare, robust and explainable evaluation of their ability to assess clinical trial reporting according to CONSORT standards remains an open challenge. In particular,…

Artificial Intelligence · Computer Science 2026-02-26 Sohyeon Jeon , Hyung-Chul Lee

As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM…

As Artificial Intelligence (AI) technologies continue to evolve, the gap between academic AI education and real-world industry challenges remains an important area of investigation. This study provides preliminary insights into challenges…

Computers and Society · Computer Science 2025-05-07 Mahir Akgun , Hadi Hosseini

Language models (LMs) are increasingly used to simulate human-like responses in scenarios where accurately mimicking a population's behavior can guide decision-making, such as in developing educational materials and designing public…

Computation and Language · Computer Science 2024-07-23 Joy He-Yueya , Wanjing Anya Ma , Kanishk Gandhi , Benjamin W. Domingue , Emma Brunskill , Noah D. Goodman

Are AI systems truly representing human values, or merely averaging across them? Our study suggests a concerning reality: Large Language Models (LLMs) fail to represent diverse cultural moral frameworks despite their linguistic…

Computation and Language · Computer Science 2025-08-01 Simon Münker

Large Language Models (LLMs) are increasingly proposed as near-autonomous artificial intelligence (AI) agents capable of making everyday decisions on behalf of humans. Although LLMs perform well on many technical tasks, their behaviour in…

Computation and Language · Computer Science 2025-06-24 Mateusz Cedro , Timour Ichmoukhamedov , Sofie Goethals , Yifan He , James Hinns , David Martens

Large language models (LLMs) have achieved strong performance on reasoning benchmarks, yet their ability to solve real-world problems requiring end-to-end workflows remains unclear. Mathematical modeling competitions provide a stringent…

Computation and Language · Computer Science 2026-04-07 Yuhang Liu , Heyan Huang , Yizhe Yang , Hongyan Zhao , Zhizhuo Zeng , Yang Gao

There has been much recent interest in evaluating large language models for uncertainty calibration to facilitate model control and modulate user trust. Inference time uncertainty, which may provide a real-time signal to the model or…

Computation and Language · Computer Science 2025-08-12 Kyle Moore , Jesse Roberts , Daryl Watson