English
Related papers

Related papers: Simultaneous Confidence Intervals for Ranks With A…

200 papers

Complex scientific models where the likelihood cannot be evaluated present a challenge for statistical inference. Over the past two decades, a wide range of algorithms have been proposed for learning parameters in computationally feasible…

Computation · Statistics 2021-12-16 Aden Forrow , Ruth E. Baker

Rankings are central to decision-making in fields ranging from education to online platforms, yet classical deterministic methods such as the Borda count method or Copeland-type pairwise methods ignore uncertainty due to sampling noise or…

Methodology · Statistics 2026-05-20 Shunpu Zhang

Increasing test-time computation is a straightforward approach to enhancing the quality of responses in Large Language Models (LLMs). While Best-of-N sampling and Self-Consistency with majority voting are simple and effective, they require…

Machine Learning · Computer Science 2025-03-04 Chengsong Huang , Langlin Huang , Jixuan Leng , Jiacheng Liu , Jiaxin Huang

This paper proposes a decorrelation-based approach to test hypotheses and construct confidence intervals for the low dimensional component of high dimensional proportional hazards models. Motivated by the geometric projection principle, we…

Machine Learning · Statistics 2014-12-17 Ethan X. Fang , Yang Ning , Han Liu

Inference scaling helps LLMs solve complex reasoning problems through extended runtime computation. On top of long chain-of-thought (long-CoT) models, purely inference-time techniques such as best-of-N (BoN) sampling, majority voting, or…

New methods for solving the college admissions problem with indifference are presented and characterised with a Monte Carlo simulation in a variety of simple scenarios. Based on a qualifier defined as the average rank, it is found that…

Computer Science and Game Theory · Computer Science 2016-09-28 Wouter Hulsbergen

We present a novel multilabel/ranking algorithm working in partial information settings. The algorithm is based on 2nd-order descent methods, and relies on upper-confidence bounds to trade-off exploration and exploitation. We analyze this…

Machine Learning · Computer Science 2013-01-17 Claudio Gentile , Francesco Orabona

Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking models in this regime remains underexplored. We formalize dense benchmark ranking under test-time scaling and introduce Scorio, a library that…

Machine Learning · Computer Science 2026-05-12 Mohsen Hariri , Michael Hinczewski , Jing Ma , Vipin Chaudhary

Accurately and efficiently estimating system performance under uncertainty is paramount in power system planning and operation. Monte Carlo simulation is often used for this purpose, but convergence may be slow, especially when detailed…

Computation · Statistics 2020-10-23 Simon Tindemans , Goran Strbac

The Cox regression model and its associated hazard ratio (HR) are frequently used for summarizing the effect of treatments on time to event outcomes. However, the HR's interpretation strongly depends on the assumed underlying survival…

Methodology · Statistics 2021-08-10 Pablo Martinez-Camblor , Todd A. MacKenzie , A. James O'Malley

Demand for high-performance, robust, and safe autonomous systems has grown substantially in recent years. These objectives motivate the desire for efficient safety-theoretic reasoning that can be embedded in core decision-making tasks such…

Robotics · Computer Science 2022-12-27 Kristoffer M. Frey , Ted J. Steiner , Jonathan P. How

In the context of the widely used competing risks set-up we discuss different inference procedures for testing equality of two cumulative incidence functions, where the data may be subject to independent right-censoring or left-truncation.…

Statistics Theory · Mathematics 2015-10-13 Dennis Dobler , Markus Pauly

To provide users insight into the value and limits of world university rankings, a comparative analysis is conducted of 5 ranking systems: ARWU, Leiden, THE, QS and U-Multirank. It links these systems with one another at the level of…

Digital Libraries · Computer Science 2016-12-06 Henk F. Moed

Global pandemics, such as the recent COVID-19 crisis, highlight the need for stochastic epidemic models that can capture the randomness inherent in the spread of disease. Such models must be accompanied by methods for estimating parameters…

Quantitative Methods · Quantitative Biology 2026-04-13 Vincent Wieland , Nils Wassmuth , Lorenzo Contento , Martin Kühn , Jan Hasenauer

The efficiency of a Markov chain Monte Carlo algorithm might be measured by the cost of generating one independent sample, or equivalently, the total cost divided by the effective sample size, defined in terms of the integrated…

Computation · Statistics 2017-05-12 Youhan Fang , Yudong Cao , Robert D. Skeel

The Leiden Rankings can be used for grouping research universities by considering universities which are not statistically significantly different as homogeneous sets. The groups and intergroup relations can be analyzed and visualized using…

Digital Libraries · Computer Science 2018-10-16 Loet Leydesdorff , Lutz Bornmann , John Mingers

This paper provides the relevant literature with a complete toolkit for conducting robust estimation and inference about the parameters of interest involved in a high-dimensional panel data framework. Specifically, (1) we allow for…

Econometrics · Economics 2025-02-13 Jiti Gao , Fei Liu , Bin Peng , Yayi Yan

Given $m$ unknown parameters with corresponding independent estimators, the Benjamini-Hochberg (BH) procedure can be used to classify the sign of parameters such that the expected proportion of erroneous directional decisions (directional…

Methodology · Statistics 2018-05-24 Asaf Weinstein , Daniel Yekutieli

Pairwise human-preference platforms such as Chatbot Arena have become central to large language model (LLM) evaluation, yet reliable task-specific ranking remains challenging. Global leaderboards mask task heterogeneity, while ranking each…

Methodology · Statistics 2026-05-29 Jiachun Li , David Simchi-Levi , Will Wei Sun

Multiple hypothesis testing is widely used to evaluate scientific studies involving statistical tests. However, for many of these tests, p-values are not available and are thus often approximated using Monte Carlo tests such as permutation…

Applications · Statistics 2018-10-17 Axel Gandy , Georg Hahn
‹ Prev 1 8 9 10 Next ›