English
Related papers

Related papers: Assessing the protection provided by misclassifica…

200 papers

We propose a novel semi-supervised learning (SSL) method that adopts selective training with pseudo labels. In our method, we generate hard pseudo-labels and also estimate their confidence, which represents how likely each pseudo-label is…

Machine Learning · Computer Science 2021-03-16 Masato Ishii

Sample surveys are widely used to obtain information about totals, means, medians, and other parameters of finite populations. In many applications, similar information is desired for subpopulations such as individuals in specific…

Methodology · Statistics 2017-05-30 Jiahua Chen , Yukun Liu

This paper describes three methods for carrying out non-asymptotic inference on partially identified parameters that are solutions to a class of optimization problems. Applications in which the optimization problems arise include estimation…

Methodology · Statistics 2022-12-02 Joel L. Horowitz , Sokbae Lee

Boosting has garnered significant interest across both machine learning and statistical communities. Traditional boosting algorithms, designed for fully observed random samples, often struggle with real-world problems, particularly with…

Machine Learning · Statistics 2026-02-19 Yuan Bian , Grace Y. Yi , Wenqing He

This paper describes privacy-preserving approaches for the statistical analysis. It describes motivations for privacy-preserving approaches for the statistical analysis of sensitive data, presents examples of use cases where such methods…

Managers, employers, policymakers, and others often seek to understand whether decisions are biased against certain groups. One popular analytic strategy is to estimate disparities after adjusting for observed covariates, typically with a…

Applications · Statistics 2024-01-29 Jongbin Jung , Sam Corbett-Davies , Johann D. Gaebler , Ravi Shroff , Sharad Goel

Researchers in the behavioral and social sciences use linear discriminant analysis (LDA) for predictions of group membership (classification) and for identifying the variables most relevant to group separation among a set of continuous…

Methodology · Statistics 2025-05-28 Ricarda Graf , Marina Zeldovich , Sarah Friedrich

When data are collected subject to a detection limit, observations below the detection limit may be considered censored. In addition, the domain of such observations may be restricted; for example, values may be required to be non-negative.…

Applications · Statistics 2020-06-30 Justin R. Williams , Hyung-Woo Kim , Catherine M. Crespi

Given a sample of size $N$, it is often useful to select a subsample of smaller size $n<N$ to be used for statistical estimation or learning. Such a data selection step is useful to reduce the requirements of data labeling and the…

Machine Learning · Statistics 2023-10-05 Germain Kolossov , Andrea Montanari , Pulkit Tandon

Studies of the effects of medical interventions increasingly take place in distributed research settings using data from multiple clinical data sources including electronic health records and administrative claims. In such settings, privacy…

Methodology · Statistics 2021-01-06 Martijn J. Schuemie , Yong Chen , David Madigan , Marc A. Suchard

Penalized logistic regression methods are frequently used to investigate the relationship between a binary outcome and a set of explanatory variables. The model performance can be assessed by measures such as the concordance statistic…

Methodology · Statistics 2021-01-20 Angelika Geroldinger , Lara Lusa , Mariana Nold , Georg Heinze

We propose a practical methodology to protect a user's private data, when he wishes to publicly release data that is correlated with his private data, in the hope of getting some utility. Our approach relies on a general statistical…

Cryptography and Security · Computer Science 2015-10-28 Salman Salamatian , Amy Zhang , Flavio du Pin Calmon , Sandilya Bhamidipati , Nadia Fawaz , Branislav Kveton , Pedro Oliveira , Nina Taft

Statistical uncertainties complicate engineering design -- confounding regulated design approaches, and degrading the performance of reliability efforts. The simplest means to tackle this uncertainty is double loop simulation; a nested…

Methodology · Statistics 2018-11-02 Zachary del Rosario , Richard W. Fenrich , Gianluca Iaccarino

In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we…

Statistics Theory · Mathematics 2019-11-12 Dave Zachariah , Petre Stoica

Quantification is the supervised learning task that consists of training predictors of the class prevalence values of sets of unlabelled data, and is of special interest when the labelled data on which the predictor has been trained and the…

Machine Learning · Computer Science 2023-10-10 Pablo González , Alejandro Moreo , Fabrizio Sebastiani

Although randomized experiments are widely regarded as the gold standard for estimating causal effects, missing data of the pretreatment covariates makes it challenging to estimate the subgroup causal effects. When the missing data…

Statistics Theory · Mathematics 2014-01-08 Peng Ding , Zhi Geng

There is an increasing focus on reducing inequalities in health outcomes in developing countries. Subnational variation is of particular interest, with geographic data used to understand the spatial risk of detrimental outcomes and to…

Methodology · Statistics 2020-12-22 Neal Marquez , Jon Wakefield

In this work we consider a problem of multi-label classification, where each instance is associated with some binary vector. Our focus is to find a classifier which minimizes false negative discoveries under constraints. Depending on the…

Statistics Theory · Mathematics 2019-03-29 Evgenii Chzhen

Recent research in differential privacy demonstrated that (sub)sampling can amplify the level of protection. For example, for $\epsilon$-differential privacy and simple random sampling with sampling rate $r$, the actual privacy guarantee is…

Applications · Statistics 2022-02-22 Jingchen Hu , Joerg Drechsler , Hang J. Kim

In this paper we consider how to evaluate survival distribution predictions with measures of discrimination. This is a non-trivial problem as discrimination measures are the most commonly used in survival analysis and yet there is no clear…

Machine Learning · Statistics 2022-03-10 Raphael Sonabend , Andreas Bender , Sebastian Vollmer