English
Related papers

Related papers: A graphical method of cumulative differences betwe…

200 papers

A quantifier is a supervised machine learning algorithm, focused on estimating the class prevalence in a dataset rather than labeling its individual observations. We introduce Continuous Sweep, a new parametric binary quantifier inspired by…

Machine Learning · Statistics 2024-10-14 Kevin Kloos , Julian D. Karch , Quinten A. Meertens , Mark de Rooij

We investigate the estimation of subgroup treatment effects with observational data. Existing propensity score matching and weighting methods are mostly developed for estimating overall treatment effect. Although the true propensity score…

Methodology · Statistics 2017-07-20 Jing Dong , Junni L Zhang , Fan Li

Entropy is a measure of heterogeneity widely used in applied sciences, often when data are collected over space. Recently, a number of approaches has been proposed to include spatial information in entropy. The aim of entropy is to…

Statistics Theory · Mathematics 2019-11-12 Linda Altieri , Daniela Cocchi , Giulia Roli

Data-driven decision support tools play an increasingly central role in decision-making across various domains. In this work, we focus on binary classification models for predicting positive-outcome scores and deciding on resource…

Machine Learning · Computer Science 2025-04-30 Simon De Vos , Jente Van Belle , Andres Algaba , Wouter Verbeke , Sam Verboven

Convergence results for averages of independent replications of counting processes are established in a $p$-variation setting and under certain assumptions. Such convergence results can be combined with functional differentiability results…

Probability · Mathematics 2019-03-12 Morten Overgaard

Classification, the process of assigning a label (or class) to an observation given its features, is a common task in many applications. Nonetheless in most real-life applications, the labels can not be fully explained by the observed…

Machine Learning · Statistics 2018-11-07 Johan Barthélemy , Morgane Dumont , Timoteo Carletti

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

Statistics Theory · Mathematics 2020-06-01 Rémy Mariétan , Stephan Morgenthaler

Cognitive biases are widespread in humans and animals alike, and can sometimes be reinforced by social interactions. One prime bias in judgment and decision-making is the human tendency to underestimate large quantities. Previous research…

Physics and Society · Physics 2022-01-12 Bertrand Jayles , Clément Sire , Ralf H. J. M Kurvers

Principal stratification is a framework for making sense of causal effects conditioned on variables that may themselves have been affected by the treatment. For instance, in an evaluation of an educational intervention, some subjects in the…

Methodology · Statistics 2026-05-05 Adam C. Sales , Kirk P. Vanacore , Erin R. Ottmar

A frequent task in exploratory data analysis consists in examining pairwise dependencies between data variables. Popular approaches include visualizing correlation or scatter plot matrices. However, both methods can be misleading. The…

Applications · Statistics 2022-04-04 Arturo Erdely , Manuel Rubio-Sanchez

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

Statistics Theory · Mathematics 2020-03-09 Rémy Mariétan , Stephan Morgenthaler

We propose a summary measure defined as the expected value of a random variable over disjoint subsets of its support that are specified by a given grid of proportions, and consider its use in a regression modeling framework. The obtained…

Statistics Theory · Mathematics 2018-10-19 Celia García-Pareja , Matteo Bottai

Forecasts for uncertain future events should be probabilistic. Probabilistic forecasts are commonly issued as prediction intervals, which provide a measure of uncertainty in the unknown outcome whilst being easier to understand and…

Methodology · Statistics 2025-08-26 Sam Allen , Julia Burnello , Johanna Ziegel

A method is discussed that allows combining sets of differential or inclusive measurements. It is assumed that at least one measurement was obtained with simultaneously fitting a set of nuisance parameters, representing sources of…

Data Analysis, Statistics and Probability · Physics 2018-01-09 Jan Kieseler

In many practical applications, such as fraud detection, credit risk modeling or medical decision making, classification models for assigning instances to a predefined set of classes are required to be both precise as well as interpretable.…

Machine Learning · Statistics 2021-09-27 Jakob Raymaekers , Wouter Verbeke , Tim Verdonck

Assessing variability according to distinct factors in data is a fundamental technique of statistics. The method commonly regarded to as analysis of variance (ANOVA) is, however, typically confined to the case where all levels of a factor…

Methodology · Statistics 2013-03-15 Steven Geinitz , Reinhard Furrer

We consider a general statistical estimation problem wherein binary labels across different observations are not independent conditioned on their feature vectors, but dependent, capturing settings where e.g. these observations are collected…

Machine Learning · Computer Science 2021-07-22 Yuval Dagan , Constantinos Daskalakis , Nishanth Dikkala , Surbhi Goel , Anthimos Vardis Kandiros

Surveys are commonly used to facilitate research in epidemiology, health, and the social and behavioral sciences. Often, these surveys are not simple random samples, and respondents are given weights reflecting their probability of…

Methodology · Statistics 2024-08-20 Adway S. Wadekar , Jerome P. Reiter

Given a high-dimensional data set we often wish to find the strongest relationships within it. A common strategy is to evaluate a measure of dependence on every variable pair and retain the highest-scoring pairs for follow-up. This strategy…

Knowledge of the binary population in stellar groupings provides important information about the outcome of the star forming process in different environments (see, e.g., Blaauw 1991, and references therein). Binarity is also a key…