English
Related papers

Related papers: Robust Kernel Hypothesis Testing under Data Corrup…

200 papers

Permutation tests are a popular choice for distinguishing distributions and testing independence, due to their exact, finite-sample control of false positives and their minimax optimality when paired with U-statistics. However, standard…

Statistics Theory · Mathematics 2025-03-26 Carles Domingo-Enrich , Raaz Dwivedi , Lester Mackey

Real data are rarely pure. Hence the past half-century has seen great interest in robust estimation algorithms that perform well even when part of the data is corrupt. However, their vast majority approach optimal accuracy only when given a…

Machine Learning · Computer Science 2022-02-14 Ayush Jain , Alon Orlitsky , Vaishakh Ravindrakumar

The paper introduces robust independence tests with non-asymptotically guaranteed significance levels for stochastic linear time-invariant systems, assuming that the observed outputs are synchronous, which means that the systems are driven…

Machine Learning · Statistics 2023-08-07 Ambrus Tamás , Dániel Ágoston Bálint , Balázs Csanád Csáji

In hypothesis testing, the phenomenon of label noise, in which hypothesis labels are switched at random, contaminates the likelihood functions. In this paper, we develop a new method to determine the decision rule when we do not have…

Information Theory · Computer Science 2014-10-28 Dennis Wei , Kush R. Varshney

To address the shortcomings of real-world datasets, robust learning algorithms have been designed to overcome arbitrary and indiscriminate data corruption. However, practical processes of gathering data may lead to patterns of data…

Machine Learning · Computer Science 2024-05-02 Lunjia Hu , Charlotte Peale , Judy Hanwen Shen

We consider finite-sample inference for a single regression coefficient in the fixed-design linear model $Y = Z\beta + bX + \varepsilon$, where $\varepsilon\in\mathbb{R}^n$ may exhibit complex dependence or heterogeneity. We develop a group…

Methodology · Statistics 2026-04-20 Zonghan Li , Hongyi Zhou , Zhiheng Zhang

We characterize the asymptotic performance of nonparametric goodness of fit testing. The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, and a test is optimal if it achieves the maximum…

Machine Learning · Statistics 2019-03-19 Shengyu Zhu , Biao Chen , Pengfei Yang , Zhitang Chen

Invariance-based randomization tests -- such as permutation tests, rotation tests, or sign changes -- are an important and widely used class of statistical methods. They allow drawing inferences under weak assumptions on the data…

Statistics Theory · Mathematics 2022-05-31 Edgar Dobriban

The last decade witnessed an explosion in the availability of data for operations research applications. Motivated by this growing availability, we propose a novel schema for utilizing data to design uncertainty sets for robust optimization…

Optimization and Control · Mathematics 2014-11-25 Dimitris Bertsimas , Vishal Gupta , Nathan Kallus

Do two data samples come from different distributions? Recent studies of this fundamental problem focused on embedding probability distributions into sufficiently rich characteristic Reproducing Kernel Hilbert Spaces (RKHSs), to compare…

Machine Learning · Computer Science 2013-05-03 Somayeh Danafar , Paola M. V. Rancoita , Tobias Glasmachers , Kevin Whittingstall , Juergen Schmidhuber

We develop a new permutation test for inference on a subvector of coefficients in linear models. The test is exact when the regressors and the error terms are independent. Then, we show that the test is asymptotically of correct level,…

Econometrics · Economics 2023-09-13 Xavier D'Haultfœuille , Purevdorj Tuvaandorj

A common challenge in nonparametric inference is its high computational complexity when data volume is large. In this paper, we develop computationally efficient nonparametric testing by employing a random projection strategy. In the…

Statistics Theory · Mathematics 2018-02-20 Meimei Liu , Zuofeng Shang , Guang Cheng

We present a general framework for hypothesis testing on distributions of sets of individual examples. Sets may represent many common data sources such as groups of observations in time series, collections of words in text or a batch of…

Methodology · Statistics 2021-02-03 Alexis Bellot , Mihaela van der Schaar

We consider the problem of constructing robust nonparametric confidence intervals and tests of hypothesis for the median when the data distribution is unknown and the data may contain a small fraction of contamination. We propose a…

Statistics Theory · Mathematics 2007-06-13 Victor J. Yohai , Ruben H. Zamar

Nonparametric tests via kernel embedding of distributions have witnessed a great deal of practical successes in recent years. However, statistical properties of these tests are largely unknown beyond consistency against a fixed alternative.…

Statistics Theory · Mathematics 2019-09-10 Tong Li , Ming Yuan

Randomly censored survival data are frequently encountered in applied sciences including biomedical or reliability applications and clinical trial analyses. Testing the significance of statistical hypotheses is crucial in such analyses to…

Methodology · Statistics 2019-01-08 Abhik Ghosh , Ayanendranath Basu , Leandro Pardo

The minimax robust hypothesis testing problem for the case where the nominal probability distributions are subject to both modeling errors and outliers is studied in twofold. First, a robust hypothesis testing scheme based on a relative…

Information Theory · Computer Science 2015-02-04 Gökhan Gül , Abdelhak M. Zoubir

Not all experiments publish their results with a description of the correlations between the data points. This makes it difficult to do hypothesis tests or model fits with that data, since just assuming no correlation can lead to an over-…

Data Analysis, Statistics and Probability · Physics 2021-06-30 Lukas Koch

We study robust mean estimation in an online and distributed scenario in the presence of adversarial data attacks. At each time step, each agent in a network receives a potentially corrupted data point, where the data points were originally…

Cryptography and Security · Computer Science 2022-09-21 Tong Yao , Shreyas Sundaram

We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However,…

Machine Learning · Statistics 2026-05-05 Gyumin Lee , Shubhanshu Shekhar , Ilmun Kim