English
Related papers

Related papers: When prompt perturbations break your A/B test: A v…

200 papers

Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy…

Computation and Language · Computer Science 2025-10-17 Jens Rupprecht , Georg Ahnert , Markus Strohmaier

Permutation tests are a distribution free way of performing hypothesis tests. These tests rely on the condition that the observed data are exchangeable among the groups being tested under the null hypothesis. This assumption is easily…

Methodology · Statistics 2017-12-14 Daniell Toth

Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results,…

Machine Learning · Computer Science 2025-02-25 Sarah Ball , Simeon Allmendinger , Frauke Kreuter , Niklas Kühl

The versatility of Large Language Models (LLMs) on natural language understanding tasks has made them popular for research in social sciences. To properly understand the properties and innate personas of LLMs, researchers have performed…

Computation and Language · Computer Science 2024-04-03 Bangzhao Shu , Lechen Zhang , Minje Choi , Lavinia Dunagan , Lajanugen Logeswaran , Moontae Lee , Dallas Card , David Jurgens

Can large language models (LLMs) simulate social surveys? To answer this question, we conducted millions of simulations in which LLMs were asked to answer subjective questions. A comparison of different LLM responses with the European…

Computation and Language · Computer Science 2024-10-22 Mingmeng Geng , Sihong He , Roberto Trotta

Given an input query, generative models such as large language models produce a random response drawn from a response distribution. Given two input queries, it is natural to ask if their response distributions are the same. While…

Statistics Theory · Mathematics 2025-09-16 Aranyak Acharyya , Carey E. Priebe , Hayden S. Helm

With the widespread adoption of pre-trained Large Language Models (LLM), there exists a high demand for task-specific test sets to benchmark their performance in domains such as healthcare and biomedicine. However, the cost of labeling test…

Computation and Language · Computer Science 2026-03-23 Aashish Anantha Ramakrishnan , Ardavan Saeedi , Hamid Reza Hassanzadeh , Fazlolah Mohaghegh , Dongwon Lee

Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model's average performance across the test prompts of a benchmark to evaluate the model's performance.…

Computation and Language · Computer Science 2024-06-07 Melissa Ailem , Katerina Marazopoulou , Charlotte Siska , James Bono

As large language models (LLMs) become more capable, there is growing excitement about the possibility of using LLMs as proxies for humans in real-world tasks where subjective labels are desired, such as in surveys and opinion polling. One…

Computation and Language · Computer Science 2024-02-07 Lindia Tjuatja , Valerie Chen , Sherry Tongshuang Wu , Ameet Talwalkar , Graham Neubig

Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whether multi-attribute persona prompting improves LLM…

Computers and Society · Computer Science 2026-02-24 Erika Elizabeth Taday Morocho , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

This study investigates the reasoning robustness of large language models (LLMs) on mathematical problem-solving tasks under systematically introduced input perturbations. Using the GSM8K dataset as a controlled testbed, we evaluate how…

Artificial Intelligence · Computer Science 2025-04-04 Giannis Chatziveroglou , Richard Yun , Maura Kelleher

Statistical significance tests can provide evidence that the observed difference in performance between two methods is not due to chance. In Information Retrieval, some studies have examined the validity and suitability of such tests for…

Information Retrieval · Computer Science 2019-04-09 Javier Parapar , David E. Losada , Manuel A. Presedo-Quindimil , Alvaro Barreiro

What counts as evidence for syntactic structure? In traditional generative grammar, systematic contrasts in grammaticality such as subject-auxiliary inversion and the licensing of parasitic gaps are taken as evidence for an internal,…

Computation and Language · Computer Science 2025-12-12 Lars G. B. Johnsen

Large language models (LLMs) have become mainstream technology with their versatile use cases and impressive performance. Despite the countless out-of-the-box applications, LLMs are still not reliable. A lot of work is being done to improve…

Computation and Language · Computer Science 2023-06-13 Aisha Khatun , Daniel G. Brown

Large Language Models (LLMs) can comply with harmful instructions, raising serious safety concerns despite their impressive capabilities. Recent work has leveraged probing-based approaches to study the separability of malicious and benign…

Computation and Language · Computer Science 2025-12-16 Cheng Wang , Zeming Wei , Qin Liu , Muhao Chen

Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models' probability distributions over strings, prompting requires models to access this…

Computation and Language · Computer Science 2023-10-24 Jennifer Hu , Roger Levy

As Large Language Models (LLMs) become increasingly embedded in empirical research workflows, their use as analytical tools for quantitative or qualitative data raises pressing concerns for scientific integrity. This opinion paper draws a…

Human-Computer Interaction · Computer Science 2025-08-12 Thomas Kosch , Sebastian Feger

Survey research is a fundamental empirical method in software engineering, enabling the systematic collection of data on professional practices, perceptions, and experiences. However, recent advances in large language models (LLMs) have…

Software Engineering · Computer Science 2025-12-22 Ronnie de Souza Santos , Italo Santos , Maria Teresa Baldassarre , Cleyton Magalhaes , Mairieli Wessel

Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large language model (LLM) performance, has been widely accepted as a…

Computation and Language · Computer Science 2025-09-03 Andong Hua , Kenan Tang , Chenhe Gu , Jindong Gu , Eric Wong , Yao Qin

Consider the problem of testing whether the outputs of a large language model (LLM) system change under an arbitrary intervention, such as an input perturbation or changing the model variant. We cannot simply compare two LLM outputs since…

Computation and Language · Computer Science 2025-06-10 Paulius Rauba , Qiyao Wei , Mihaela van der Schaar
‹ Prev 1 2 3 10 Next ›