中文
相关论文

相关论文: Sets of Half-Average Nulls Generate Risk-Limiting …

200 篇论文

Large Language Models (LLMs) offer a promising alternative to traditional survey methods, potentially enhancing efficiency and reducing costs. In this study, we use LLMs to create virtual populations that answer survey questions, enabling…

人机交互 · 计算机科学 2025-03-24 Enzo Sinacola , Arnault Pachot , Thierry Petit

Sample-average approximations (SAA) are a practical means of finding approximate solutions of stochastic programming problems involving an extremely large (or infinite) number of scenarios. SAA can also be used to find estimates of a lower…

其他统计学 · 统计学 2014-05-08 Jiajie Chen , Cong Han Lim , Peter Z. G. Qian , Jeff Linderoth , Stephen J. Wright

This paper presents a systematic analysis of biases in open-source Large Language Models (LLMs), across gender, religion, and race. Our study evaluates bias in smaller-scale Llama and Gemma models using the SALT ($\textbf{S}$ocial…

计算与语言 · 计算机科学 2025-02-19 Samee Arif , Zohaib Khan , Maaidah Kaleem , Suhaib Rashid , Agha Ali Raza , Awais Athar

Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We conceptualise auditing as…

机器学习 · 计算机科学 2026-01-07 David Hartmann , Lena Pohlmann , Lelia Hanslik , Noah Gießing , Bettina Berendt , Pieter Delobelle

Reasoning-capable LLMs have achieved major breakthroughs in solving complex problems, but recent work shows that acquiring and deploying strong reasoning can introduce significant safety risks. A common mitigation is to apply a secondary…

人工智能 · 计算机科学 2026-02-03 Yihao Xue , Baharan Mirzasoleiman

We study the relative-error property testing model for Boolean functions that was recently introduced in the work of Chen et al. (SODA 2025). In relative-error testing, the testing algorithm gets uniform random satisfying assignments as…

计算复杂性 · 计算机科学 2025-04-15 Xi Chen , William Pires , Toniann Pitassi , Rocco A. Servedio

Roboticists usually test new control software in simulation environments before evaluating its functionality on real-world robots. Simulations reduce the risk of damaging the hardware and can significantly increase the development process's…

机器人学 · 计算机科学 2021-02-17 Felix Sygulla , Daniel Rixen

Masked Language Models (MLMs) have shown superior performances in numerous downstream NLP tasks when used as text encoders. Unfortunately, MLMs also demonstrate significantly worrying levels of social biases. We show that the previously…

计算与语言 · 计算机科学 2021-04-16 Masahiro Kaneko , Danushka Bollegala

One of the main challenges developers face in the use of continuous integration (CI) and deployment pipelines is the occurrence of intermittent job failures, which result from unexpected non-deterministic issues (e.g., flaky tests or…

软件工程 · 计算机科学 2026-01-21 Henri Aïdasso , Francis Bordeleau , Ali Tizghadam

Metamorphic testing (MT) alleviates the oracle problem by checking metamorphic relations (MRs) across multiple test executions. The fault detection effectiveness of MT is influenced not only by the choice and quality of MRs, but also by how…

软件工程 · 计算机科学 2025-12-16 Zenghui Zhou , Pak-Lok Poon , Zheng Zheng , Xiao-Yi Zhang

Noisy labels are common in large-scale medical imaging datasets due to inter-observer variability and ambiguous cases. We propose a statistically grounded and task-agnostic framework, Standardized Loss Aggregation (SLA), for detecting noisy…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Inhyuk Park , Doohyun Park

In this paper we develop a novel approach to relaxing Arrow's axioms for voting rules, addressing a long-standing critique in social choice theory. Classical axioms (often styled as fairness axioms or fairness criteria) are assessed in a…

计算机科学与博弈论 · 计算机科学 2025-06-17 Suvadip Sana , Daniel Brous , Martin T. Wells , Moon Duchin

Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on artificial prompts that poorly reflect real-world use or on naturalistic tasks that lack scale and…

计算与语言 · 计算机科学 2026-05-12 Akram Elbouanani , Aboubacar Tuo , Adrian Popescu

Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on hard ones. We adapt Wald's Sequential Probability Ratio Test (SPRT) as a plug-in compute…

机器学习 · 计算机科学 2026-05-20 Andrea Morandi

We introduce ZeroSCROLLS, a zero-shot benchmark for natural language understanding over long texts, which contains only test and small validation sets, without training data. We adapt six tasks from the SCROLLS benchmark, and add four new…

计算与语言 · 计算机科学 2023-12-19 Uri Shaham , Maor Ivgi , Avia Efrat , Jonathan Berant , Omer Levy

Observed events in recommendation are consequence of the decisions made by a policy, thus they are usually selectively labeled, namely the data are Missing Not At Random (MNAR), which often causes large bias to the estimate of true outcomes…

机器学习 · 计算机科学 2021-10-05 Zifeng Wang , Xi Chen , Rui Wen , Shao-Lun Huang

Online A/B testing plays a critical role in the high-tech industry to guide product development and accelerate innovation. It performs a null hypothesis statistical test to determine which variant is better. However, a typical A/B test…

统计方法学 · 统计学 2021-09-03 Miao Yu , Wenbin Lu , Rui Song

Ranking LLMs via pairwise human feedback underpins current leaderboards for open-ended tasks, such as creative writing and problem-solving. We analyze ~89K comparisons in 116 languages from 52 LLMs from Arena, and show that the best-fit…

机器学习 · 计算机科学 2026-05-08 Jai Moondra , Ayela Chughtai , Bhargavi Lanka , Swati Gupta

Representational similarity analysis (RSA) is a multivariate technique to investigate cortical representations of objects or constructs. While avoiding ill-posed matrix inversions that plague multivariate approaches in the presence of many…

统计方法学 · 统计学 2021-12-03 Roberto Viviani

Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language…