English
Related papers

Related papers: An exact, unconditional, nuisance-agnostic test fo…

200 papers

In survival studies, classical inferences for left-truncated data require quasi-independence, a property that the joint density of truncation time and failure time is factorizable into their marginal densities in the observable region. The…

Methodology · Statistics 2019-04-16 Young-Geun Choi , Wei-Yann Tsai , Myunghee Cho Paik

The lack of non-parametric statistical tests for confounding bias significantly hampers the development of robust, valid and generalizable predictive models in many fields of research. Here I propose the partial and full confounder tests,…

Machine Learning · Computer Science 2025-05-30 Tamas Spisak

When applied to high-dimensional datasets, feature selection algorithms might still leave dozens of irrelevant variables in the dataset. Therefore, even after feature selection has been applied, classifiers must be prepared to the presence…

Machine Learning · Computer Science 2018-11-21 Danilo Vasconcellos Vargas , Hirotaka Takano , Junichi Murata

We treat the problem of testing independence between m continuous variables when m can be larger than the available sample size n. We consider three types of test statistics that are constructed as sums or sums of squares of pairwise rank…

Statistics Theory · Mathematics 2016-12-05 Dennis Leung , Mathias Drton

We present methodology for constructing exact significance tests for cross tabulated data for "difficult" composite alternative hypotheses that have no natural test statistic. We construct a test for discovering Simpson's Paradox and a…

Methodology · Statistics 2013-12-12 Daniel Yekutieli

Recently, there is growing concern that machine-learning models, which currently assist or even automate decision making, reproduce, and in the worst case reinforce, bias of the training data. The development of tools and techniques for…

Programming Languages · Computer Science 2020-05-07 Caterina Urban , Maria Christakis , Valentin Wüstholz , Fuyuan Zhang

Various measures in two-way contingency table analysis have been proposed to express the strength of association between row and column variables in contingency tables. Tomizawa et al. (2004) proposed more general measures, including…

Methodology · Statistics 2023-07-10 Wataru Urasaki , Tomoyuki Nakagawa , Tomotaka Momozaki , Sadao Tomizawa

In a Monte-Carlo test, the observed dataset is fixed, and several resampled or permuted versions of the dataset are generated in order to test a null hypothesis that the original dataset is exchangeable with the resampled/permuted ones.…

Methodology · Statistics 2025-05-05 Lasse Fischer , Aaditya Ramdas

This paper provides a statistical method to test whether a system that performs a binary sequential hypothesis test is optimal in the sense of minimizing the average decision times while taking decisions with given reliabilities. The…

Information Theory · Computer Science 2018-01-08 Meik Dörpinghaus , Izaak Neri , Édgar Roldán , Heinrich Meyr , Frank Jülicher

This paper proposes new nonparametric diagnostic tools to assess the asymptotic validity of different treatment effects estimators that rely on the correct specification of the propensity score. We derive a particular restriction relating…

Methodology · Statistics 2019-02-11 Pedro H. C. Sant'Anna , Xiaojun Song

Conditional-independence-based discovery uses statistical tests to identify a graphical model that represents the independence structure of variables in a dataset. These tests, however, can be unreliable, and algorithms are sensitive to…

Machine Learning · Computer Science 2026-04-21 Philipp M. Faller , Dominik Janzing

A wide range of interesting program properties are intrinsically relational, i.e., they relate two or more program traces. Two prominent relational properties are secure information flow and conditional program equivalence. By showing the…

Logic in Computer Science · Computer Science 2019-10-22 Alexander Weigl , Mattias Ulbrich , Suhyun Cha , Bernhard Beckert , Birgit Vogel-Heuser

We consider the conditional randomization test as a way to account for covariate imbalance in randomized experiments. The test accounts for covariate imbalance by comparing the observed test statistic to the null distribution of the test…

Modern machine learning models are highly expressive but notoriously difficult to analyze statistically. In particular, while black-box predictors can achieve strong empirical performance, they rarely provide valid hypothesis tests or…

Machine Learning · Computer Science 2026-03-10 Mohamed Salem

We investigate one/two-sample mean tests for high-dimensional compositional data when the number of variables is comparable with the sample size, as commonly encountered in microbiome research. Existing methods mainly focus on max-type test…

Statistics Theory · Mathematics 2024-04-15 Qianqian Jiang , Wenbo Li , Zeng Li

We introduce a new procedure for testing the significance of a set of regression coefficients in a Gaussian linear model with $n \geq d$. Our method, the $L$-test, provides the same statistical validity guarantee as the classical $F$-test,…

Methodology · Statistics 2025-12-01 Danielle Paulson , Souhardya Sengupta , Lucas Janson

Causal inference in completely randomized treatment-control studies with binary outcomes is discussed from Fisherian, Neymanian and Bayesian perspectives, using the potential outcomes framework. A randomization-based justification of…

Statistics Theory · Mathematics 2015-01-13 Peng Ding , Tirthankar Dasgupta

In this paper, we investigate score function-based tests to check the significance of an ultrahigh-dimensional sub-vector of the model coefficients when the nuisance parameter vector is also ultrahigh-dimensional in linear models. We first…

Methodology · Statistics 2024-11-12 Weichao Yang , Xu Guo , Lixing Zhu

In modern scientific research, small-scale studies with limited participants are increasingly common. However, interpreting individual outcomes can be challenging, making it standard practice to combine data across studies using random…

Statistics Theory · Mathematics 2025-11-04 Lucas Kania , Larry Wasserman , Sivaraman Balakrishnan

We construct examples of contingency tables on $n$ binary random variables where the gap between the linear programming lower/upper bound and the true integer lower/upper bounds on cell entries is exponentially large. These examples provide…

Optimization and Control · Mathematics 2007-06-13 Seth Sullivant