English
Related papers

Related papers: Predictively Consistent Prior Effective Sample Siz…

200 papers

Discrete entropy estimation is a classic information theory problem, wherein the average information content of a discrete random variable is estimated from samples alone. Naive approaches, such as the plugin method, fail to account for the…

Information Theory · Computer Science 2026-05-04 Lucas H. McCabe , H. Howie Huang

This paper is concerned with sample size determination methodology for prediction models. We propose combining the individual calculations via a learning-type curve. We suggest two distinct ways of doing so, a deterministic skeleton of a…

Methodology · Statistics 2024-05-24 Alimu Dayimu , Nikola Simidjievski , Nikolaos Demiris , Jean Abraham

Event Sequences (EvS) refer to sequential data characterized by irregular sampling intervals and a mix of categorical and numerical features. Accurate classification of these sequences is crucial for various real-life applications,…

Machine Learning · Computer Science 2025-02-27 Dmitry Osin , Igor Udovichenko , Viktor Moskvoretskii , Egor Shvetsov , Evgeny Burnaev

When considering a model selection or, more generally, an aggregation approach for adaptive statistical inference, it is often necessary to compute estimators over a wide range of model complexities including unnecessarily large models even…

Statistics Theory · Mathematics 2026-04-17 Ilsang Ohn , Shitao Fan , Jungbin Jun , Lizhen Lin

Classical statistical methods have theoretical justification when the sample size is predetermined. In applications, however, it's often the case that sample sizes are data-dependent rather than predetermined. The aforementioned methods…

Statistics Theory · Mathematics 2026-05-06 Ryan Martin

A popular setting in medical statistics is a group sequential trial with independent and identically distributed normal outcomes, in which interim analyses of the sum of the outcomes are performed. Based on a prescribed stopping rule, one…

Statistics Theory · Mathematics 2018-04-03 Ben Berckmoes , Anna Ivanova , Geert Molenberghs

A method for large scale Gaussian process classification has been recently proposed based on expectation propagation (EP). Such a method allows Gaussian process classifiers to be trained on very large datasets that were out of the reach of…

We introduce a Bayesian prior distribution, the Logit-Normal continuous analogue of the spike-and-slab (LN-CASS), which enables flexible parameter estimation and variable/model selection in a variety of settings. We demonstrate its use and…

Applications · Statistics 2018-10-04 William Thomson , Sara Jabbari , Angela Taylor , Wiebke Arlt , David Smith

We study the effects of data size and quality on the performance on Automated Essay Scoring (AES) engines that are designed in accordance with three different paradigms; A frequency and hand-crafted feature-based model, a recurrent neural…

Computation and Language · Computer Science 2021-08-31 Christopher Ormerod , Amir Jafari , Susan Lottridge , Milan Patel , Amy Harris , Paul van Wamelen

We propose the use of a simple intuitive principle for measuring algorithmic classification bias: the significance of the differences in a classifier's error rates across the various demographics is inversely commensurate with the sample…

Methodology · Statistics 2026-01-08 Ioannis Ivrissimtzis , Shauna Concannon , Matthew Houliston , Graham Roberts

Predictive posterior densities (PPDs) are of interest in approximate Bayesian inference. Typically, these are estimated by simple Monte Carlo (MC) averages using samples from the approximate posterior. We observe that the signal-to-noise…

Machine Learning · Computer Science 2024-05-31 Abhinav Agrawal , Justin Domke

We consider the linear regression problem under semi-supervised settings wherein the available data typically consists of: (i) a small or moderate sized 'labeled' data, and (ii) a much larger sized 'unlabeled' data. Such data arises…

Methodology · Statistics 2018-07-02 Abhishek Chakrabortty , Tianxi Cai

Learning joint probability distributions on n random variables requires exponential sample size in the generic case. Here we consider the case that a temporal (or causal) order of the variables is known and that the (unknown) graph of…

Machine Learning · Computer Science 2007-05-23 Pawel Wocjan , Dominik Janzing , Thomas Beth

Background: Pairwise and network meta-analyses using fixed effect and random effects models are commonly applied to synthesise evidence from randomised controlled trials. The models differ in their assumptions and the interpretation of the…

Methodology · Statistics 2017-08-04 Shijie Ren , Jeremy E. Oakley , John W. Stevens

Prediction sets provide a means of quantifying the uncertainty in predictive tasks. Using held out calibration data, conformal prediction and risk control can produce prediction sets that exhibit statistically valid error control in a…

Machine Learning · Statistics 2026-02-05 Bror Hultberg , Dave Zachariah , Antônio H. Ribeiro

A new framework of compressive sensing (CS), namely statistical compressive sensing (SCS), that aims at efficiently sampling a collection of signals that follow a statistical distribution and achieving accurate reconstruction on average, is…

Computer Vision and Pattern Recognition · Computer Science 2010-10-22 Guoshen Yu , Guillermo Sapiro

Probabilistic values, including Shapley values and semivalues, provide a model-agnostic framework to attribute the behavior of a black-box model to data points or features, with a wide range of applications including explainable artificial…

Artificial Intelligence · Computer Science 2026-05-05 Ziqi Liu , Kiljae Lee , Yuan Zhang , Weijing Tang

This paper outlines a framework for quantifying the prior's contribution to posterior inference in the presence of prior-likelihood discordance, a broader concept than the usual notion of prior-likelihood conflict. We achieve this dual…

Methodology · Statistics 2021-01-08 Matthew Reimherr , Xiao-Li Meng , Dan L. Nicolae

The proposed approach extends the confidence posterior distribution to the semi-parametric empirical Bayes setting. Whereas the Bayesian posterior is defined in terms of a prior distribution conditional on the observed data, the confidence…

Methodology · Statistics 2012-05-02 David R. Bickel

The Expected Calibration Error (ece), the dominant calibration metric in machine learning, compares predicted probabilities against empirical frequencies of binary outcomes. This is appropriate when labels are binary events. However, many…

Machine Learning · Computer Science 2026-03-17 Michael Leznik