English
Related papers

Related papers: Data Reuse and the Long Shadow of Error: Splitting…

200 papers

Detecting whether examples belong to a given in-distribution or are Out-Of-Distribution (OOD) requires identifying features specific to the in-distribution. In the absence of labels, these features can be learned by self-supervised…

Artificial Intelligence · Computer Science 2022-01-19 Nima Rafiee , Rahil Gholamipoorfard , Nikolas Adaloglou , Simon Jaxy , Julius Ramakers , Markus Kollmann

Interval-censored competing risks data arise when each study subject may experience an event or failure from one of several causes and the failure time is not observed exactly but rather known to lie in an interval between two successive…

Methodology · Statistics 2016-03-02 Lu Mao , D. Y. Lin , Donglin Zeng

The design and complexity analysis of randomized coordinate descent methods, and in particular of variants which update a random subset (sampling) of coordinates in each iteration, depends on the notion of expected separable…

Optimization and Control · Mathematics 2015-05-29 Zheng Qu , Peter Richtárik

Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard…

Methodology · Statistics 2026-05-14 Marcell T. Kurbucz

We extend Fisher's randomization test (FRT) to test conditional independence between observed outcomes and treatments given covariates in both randomized experiments and observational studies, with no restriction on the variable type of…

Methodology · Statistics 2025-06-12 Zhen Zhong

Multiple testing is a fundamental problem in high-dimensional statistical inference. Although many methods have been proposed to control false discoveries, it is still a challenging task when the tests are correlated to each other. To…

Statistics Theory · Mathematics 2022-07-06 Meng Mei , Yuan Jiang

Regression adjustment is broadly applied in randomized trials under the premise that it usually improves the precision of a treatment effect estimator. However, previous work has shown that this is not always true. To further understand…

Methodology · Statistics 2022-10-11 Katarzyna Reluga , Ting Ye , Qingyuan Zhao

Randomized experiments are the gold standard for causal inference but face significant challenges in business applications, including limited traffic allocation, the need for heterogeneous treatment effect estimation, and the complexity of…

Methodology · Statistics 2025-08-18 Zhenkang Peng , Chengzhang Li , Ying Rong , Renyu Zhang

An informative sampling design leads to the selection of units whose inclusion probabilities are correlated with the response variable of interest. Model inference performed on the resulting observed sample will be biased for the population…

Methodology · Statistics 2018-06-29 Matthew R. Williams , Terrance D. Savitsky

Capturing the dependence structure of multivariate extreme events is a major concern in many fields involving the management of risks stemming from multiple sources, e.g. portfolio monitoring, insurance, environmental risk management and…

Machine Learning · Statistics 2016-03-15 Nicolas Goix , Anne Sabourin , Stéphan Clémençon

In fitting a mixture of linear regression models, normal assumption is traditionally used to model the error and then regression parameters are estimated by the maximum likelihood estimators (MLE). This procedure is not valid if the normal…

Methodology · Statistics 2018-11-06 Yanyuan Ma , Shaoli Wang , Lin Xu , Weixin Yao

When performing supervised learning with the model selected using validation error from sample splitting and cross validation, the minimum value of the validation error can be biased downward. We propose two simple methods that use the…

Methodology · Statistics 2018-02-13 Leying Guan

We study the problem of testing discrete distributions with a focus on the high probability regime. Specifically, given samples from one or more discrete distributions, a property $\mathcal{P}$, and parameters $0< \epsilon, \delta <1$, we…

Data Structures and Algorithms · Computer Science 2020-09-15 Ilias Diakonikolas , Themis Gouleakis , Daniel M. Kane , John Peebles , Eric Price

Real-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models. Accordingly, a key question for any supervised or unsupervised model is…

Machine Learning · Computer Science 2023-10-03 Zhilin Zhao , Longbing Cao

Within the machine learning community, the widely-used uniform convergence framework has been used to answer the question of how complex, over-parameterized models can generalize well to new data. This approach bounds the test error of the…

Machine Learning · Statistics 2021-03-05 Ryan Theisen , Jason M. Klusowski , Michael W. Mahoney

Inferring causal directions on discrete and categorical data is an important yet challenging problem. Even though the additive noise models (ANMs) approach can be adapted to the discrete data, the functional structure assumptions make it…

Machine Learning · Statistics 2021-09-02 Austin Goddard , Yu Xiang

Confounding is a significant obstacle to unbiased estimation of causal effects from observational data. For settings with high-dimensional covariates -- such as text data, genomics, or the behavioral social sciences -- researchers have…

Artificial Intelligence · Computer Science 2024-02-01 Katherine A. Keith , Sergey Feldman , David Jurgens , Jonathan Bragg , Rohit Bhattacharya

Finding interdependency relations between (possibly multivariate) time series provides valuable knowledge about the processes that generate the signals. Information theory sets a natural framework for non-parametric measures of several…

Information Theory · Computer Science 2016-02-09 German Gomez-Herrero , Wei Wu , Kalle Rutanen , Miguel C. Soriano , Gordon Pipa , Raul Vicente

Large-scale multiple testing is a fundamental problem in high dimensional statistical inference. It is increasingly common that various types of auxiliary information, reflecting the structural relationship among the hypotheses, are…

Methodology · Statistics 2021-10-07 Hongyuan Cao , Jun Chen , Xianyang Zhang

Statisticians increasingly face the problem to reconsider the adaptability of classical inference techniques. In particular, divers types of high-dimensional data structures are observed in various research areas; disclosing the boundaries…

Statistics Theory · Mathematics 2017-06-09 Paavo Sattler , Markus Pauly