English
Related papers

Related papers: Concordance probability in a big data setting: app…

200 papers

We propose new summary measures of diagnostic test accuracy which can be used as companions to existing diagnostic accuracy measures. Conceptually, our summary measures are tantamount to the so-called Hellinger affinity and we show that…

Methodology · Statistics 2017-12-29 Miguel de Carvalho , Bradley J. Barney , Garritt L. Page

Popular measures of meta-analysis heterogeneity, such as $I^2$, cannot be considered measures of population heterogeneity since they are dependant on samples sizes within studies. The coefficient of variation (CV) recently introduced and…

Methodology · Statistics 2020-10-06 Maxwell Cairns , Luke Prendergast

Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of information for example-based explanations. As humans cannot interpret thousands of documents,…

Computation and Language · Computer Science 2026-04-10 Loris Schoenegger , Benjamin Roth

Consider a finite sample from an unknown distribution over a countable alphabet. Unobserved events are alphabet symbols which do not appear in the sample. Estimating the probabilities of unobserved events is a basic problem in statistics…

Statistics Theory · Mathematics 2022-11-08 Amichai Painsky

It has been shown that for the analysis of X-ray spectra the C-statistic, contrary to the chi^2-statistic, provides unbiased estimates of the model parameters and their uncertainty ranges. However, it is often stated that the C-statistic…

High Energy Astrophysical Phenomena · Physics 2017-09-13 J. S. Kaastra

We consider the standard non-parametric regression model with Gaussian errors but where the data consist of different samples. The question to be answered is whether the samples can be adequately represented by the same regression function.…

Statistics Theory · Mathematics 2008-09-17 A. Kovac , P. L. Davies

The learning of domain-invariant representations in the context of domain adaptation with neural networks is considered. We propose a new regularization method that minimizes the discrepancy between domain-specific latent feature…

Approximate Bayesian Computation (ABC) is a powerful method for carrying out Bayesian inference when the likelihood is computationally intractable. However, a drawback of ABC is that it is an approximate method that induces a systematic…

Methodology · Statistics 2015-09-29 Minh Ngoc Tran , Robert Kohn

We study the frequentist properties of confidence intervals computed by the method known to statisticians as the Profile Likelihood. It is seen that the coverage of these intervals is surprisingly good over a wide range of possible…

Data Analysis, Statistics and Probability · Physics 2009-11-10 Wolfgang A. Rolke , Angel M. Lopez , Jan Conrad

Stratifying factors, like age and gender, can modify the effect of treatments and exposures on risk of a studied outcome. Several effect measures, including the relative risk, hazard ratio, odds ratio, and risk difference, can be used to…

Methodology · Statistics 2021-11-05 Jake Shannin , Babette A. Brumback

In this pedagogical text aimed at those wanting to start thinking about or brush up on probabilistic inference, I review the rules by which probability distribution functions can (and cannot) be combined. I connect these rules to the…

Data Analysis, Statistics and Probability · Physics 2012-05-22 David W. Hogg

Matching is a popular nonparametric covariate adjustment strategy in empirical health services research. Matching helps construct two groups comparable in many baseline covariates but different in some key aspects under investigation. In…

Applications · Statistics 2023-08-17 Chang Chen , Zhiyu Qian , Bo Zhang

Bayesian model comparison (BMC) offers a principled probabilistic approach to study and rank competing models. In standard BMC, we construct a discrete probability distribution over the set of possible models, conditional on the observed…

Machine Learning · Statistics 2023-02-22 Marvin Schmitt , Stefan T. Radev , Paul-Christian Bürkner

Observational cohort studies with oversampled exposed subjects are typically implemented to understand the causal effect of a rare exposure. Because the distribution of exposed subjects in the sample differs from the source population,…

Methodology · Statistics 2019-02-14 Sherri Rose

This paper introduces the concept of {mutual consensus} as a novel non-compensatory consensus measure that accounts for the maximum disparity among opinions to ensure robust consensus evaluation. Incorporating this concept, several new…

Optimization and Control · Mathematics 2025-11-04 Diego García-Zamora , Bapi Dutta , Luis Martínez

Contrastive divergence (CD) is a promising method of inference in high dimensional distributions with intractable normalizing constants, however, the theoretical foundations justifying its use are somewhat shaky. This document proposes a…

Machine Learning · Statistics 2014-05-06 Ian E Fellows

In many empirical studies of a large two-sided matching market (such as in a college admissions problem), the researcher performs statistical inference under the assumption that they observe a random sample from a large matching market. In…

Econometrics · Economics 2024-04-02 Jacob Schwartz , Kyungchul Song

The task of calibration is to retrospectively adjust the outputs from a machine learning model to provide better probability estimates on the target variable. While calibration has been investigated thoroughly in classification, it has not…

Machine Learning · Statistics 2018-06-21 Hao Song , Meelis Kull , Peter Flach

Process capability indices such as $C_{pk}$ are widely used in manufacturing quality control to support supplier qualification and product release decisions based on fixed acceptance thresholds (e.g., $C_{pk} \geq 1.33$). In practice, these…

Applications · Statistics 2026-03-13 Fei Jiang , Lei Yang

Post-click conversion rate (CVR) is a reliable indicator of online customers' preferences, making it crucial for developing recommender systems. A major challenge in predicting CVR is severe selection bias, arising from users' inherent…

Artificial Intelligence · Computer Science 2025-12-02 Wenbo Hu , Xin Sun , Qiang liu , Le Wu , Liang Wang