中文
相关论文

相关论文: Inference at Scale Significance Testing for Large …

200 篇论文

In multiple testing scenarios, typically the sign of a parameter is inferred when its estimate exceeds some significance threshold in absolute value. Typically, the significance threshold is chosen to control the experimentwise type I error…

统计方法学 · 统计学 2018-01-03 Chaoyu Yu , Peter D. Hoff

This paper reviews recent advancements in the application of optimal transport (OT) to multivariate distribution-free nonparametric testing. Inspired by classical rank-based methods, such as Wilcoxon's rank-sum and signed-rank tests, we…

统计方法学 · 统计学 2025-03-18 Zhen Huang , Bodhisattva Sen

Large language models (LLMs) struggle with multi-step reasoning, where inference-time scaling has emerged as a promising strategy for performance improvement. Verifier-guided search outperforms repeated sampling when sample size is limited…

计算与语言 · 计算机科学 2025-02-04 Fei Yu , Yingru Li , Benyou Wang

Research often necessitates of samples, yet obtaining large enough samples is not always possible. When it is, the researcher may use one of two methods for deciding upon the required sample size: rules-of-thumb, quick yet uncertain, and…

统计方法学 · 统计学 2016-04-08 Jose D. Perezgonzalez

Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of information for example-based explanations. As humans cannot interpret thousands of documents,…

计算与语言 · 计算机科学 2026-04-10 Loris Schoenegger , Benjamin Roth

While the evaluation of explanations is an important step towards trustworthy models, it needs to be done carefully, and the employed metrics need to be well-understood. Specifically model randomization testing is often overestimated and…

Feature attribution a.k.a. input salience methods which assign an importance score to a feature are abundant but may produce surprisingly different results for the same model on the same input. While differences are expected if disparate…

计算与语言 · 计算机科学 2022-11-10 Jasmijn Bastings , Sebastian Ebert , Polina Zablotskaia , Anders Sandholm , Katja Filippova

Standard tests of the "no-treatment-effect" hypothesis for a comparative experiment include permutation tests, the Wilcoxon rank sum test, two-sample $t$ tests, and Fisher-type randomization tests. Practitioners are aware that these…

统计方法学 · 统计学 2015-09-11 Joseph B. Lang

Experimental research on behavior and cognition frequently rests on stimulus or subject selection where not all characteristics can be fully controlled, even when attempting strict matching. For example, when contrasting patients to…

统计方法学 · 统计学 2016-08-29 Jona Sassenhagen , Phillip M. Alday

The evaluation of Information Retrieval (IR) systems typically uses query-document pairs with corresponding human-labelled relevance assessments (qrels). These qrels are used to determine if one system is better than another based on…

信息检索 · 计算机科学 2025-07-11 Jack McKechnie , Graham McDonald , Craig Macdonald

There are many different proposed procedures for sample size planning for the Wilcoxon-Mann-Whitney test at given type-I and type-II error rates $\alpha$ and $\beta$, respectively. Most methods assume very specific models or types of data…

统计方法学 · 统计学 2018-10-10 Martin Happ , Arne C. Bathke , Edgar Brunner

Diagnostic accuracy studies assess sensitivity and specificity of a new index test in relation to an established comparator or the reference standard. The development and selection of the index test is usually assumed to be conducted prior…

统计方法学 · 统计学 2022-08-30 Max Westphal , Antonia Zapf

The controversy about statistical significance vs. scientific relevance is more than 100 years old. But still nowadays null hypothesis significance testing is considered as gold standard in many empirical fields from economics and social…

应用统计 · 统计学 2022-11-23 Uwe Hassler

Simulation-based inference plays a major role in modern statistics, and often employs either reallocating (as in a randomization test) or resampling (as in bootstrapping). Reallocating mimics random allocation to treatment groups, while…

统计理论 · 数学 2017-08-08 Kari Lock Morgan

Given the vast reservoirs of data stored worldwide, efficient mining of data from a large information store has emerged as a great challenge. Many databases like that of intrusion detection systems, web-click records, player statistics,…

数据库 · 计算机科学 2010-03-09 Sourav Dutta , Arnab Bhattacharya

There is intense interest in investigating how inference time compute (ITC) (e.g. repeated sampling, refinements, etc) can improve large language model (LLM) capabilities. At the same time, recent breakthroughs in reasoning models, such as…

In response to growing concern about the reliability and reproducibility of published science, researchers have proposed adopting measures of greater statistical stringency, including suggestions to require larger sample sizes and to lower…

统计方法学 · 统计学 2018-07-09 Harlan Campbell , Paul Gustafson

This paper offers a commentary on the use of notions of statistical significance in choice modelling. We review the reasons for uncertainty in parameter estimates, provide a precise discussion on the computation of measures of uncertainty…

计量经济学 · 经济学 2026-05-18 Stephane Hess , Andrew Daly , Michiel Bliemer , Angelo Guevara , Ricardo Daziano , Thijs Dekker

Relevance and fairness are two major objectives of recommender systems (RSs). Recent work proposes measures of RS fairness that are either independent from relevance (fairness-only) or conditioned on relevance (joint measures). While…

信息检索 · 计算机科学 2024-05-29 Theresia Veronika Rampisela , Tuukka Ruotsalo , Maria Maistro , Christina Lioma

Recent studies on scaling up ranking models have achieved substantial improvement for recommendation systems and search engines. However, most large-scale ranking systems rely on item IDs, where each item is treated as an independent…

信息检索 · 计算机科学 2026-02-02 Zhen Zhao , Tong Zhang , Jie Xu , Qingliang Cai , Qile Zhang , Leyuan Yang , Daorui Xiao , Xiaojia Chang