English
Related papers

Related papers: Differentially private scale testing via rank tran…

200 papers

Testing mutual independence among multiple random variables is a fundamental problem in statistics, with wide applications in genomics, finance, and neuroscience. In this paper, we propose a new class of tests for high-dimensional mutual…

Applications · Statistics 2026-01-28 Ping Zhao , Huifang Ma

Background: Synthetic data has been proposed as a solution for sharing anonymized versions of sensitive biomedical datasets. Ideally, synthetic data should preserve the structure and statistical properties of the original data, while…

Machine Learning · Computer Science 2024-10-24 Ileana Montoya Perez , Parisa Movahedi , Valtteri Nieminen , Antti Airola , Tapio Pahikkala

We present the $U$-Statistic Permutation (USP) test of independence in the context of discrete data displayed in a contingency table. Either Pearson's chi-squared test of independence, or the $G$-test, are typically used for this task, but…

Methodology · Statistics 2022-01-19 Thomas B. Berrett , Richard J. Samworth

In this paper, we propose a general framework for distribution-free nonparametric testing in multi-dimensions, based on a notion of multivariate ranks defined using the theory of measure transportation. Unlike other existing proposals in…

Statistics Theory · Mathematics 2019-10-08 Nabarun Deb , Bodhisattva Sen

Learning effective data representations has been crucial in non-parametric two-sample testing. Common approaches will first split data into training and test sets and then learn data representations purely on the training set. However,…

Machine Learning · Computer Science 2025-05-09 Xunye Tian , Liuhua Peng , Zhijian Zhou , Mingming Gong , Arthur Gretton , Feng Liu

Ranking institutions such as medical centers or universities is based on an indicator accompanied with an uncertainty measure such as a standard deviation, and confidence intervals should be calculated to assess the quality of these ranks.…

Methodology · Statistics 2017-08-10 Diaa Al Mohamad , Erik W. van Zwet , Jelle J. Goeman , Aldo Solari

What proportion of treated units actually benefited from an experimental intervention? What is the median or the largest individual treatment effect? This paper develops methods for answering such questions about the distribution of…

Methodology · Statistics 2026-05-11 David Kim , Yongchang Su , Jake Bowers , Xinran Li

This work addresses testing the independence of two continuous and finite-dimensional random variables from the design of a data-driven partition. The empirical log-likelihood statistic is adopted to approximate the sufficient statistics of…

Machine Learning · Statistics 2022-01-19 Mauricio E. Gonzalez , Jorge F. Silva , Miguel Videla , Marcos E. Orchard

Differential privacy has emerged as an significant cornerstone in the realm of scientific hypothesis testing utilizing confidential data. In reporting scientific discoveries, Bayesian tests are widely adopted since they effectively…

Machine Learning · Statistics 2025-12-22 Abhisek Chakraborty , Saptati Datta

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

Applications · Statistics 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

We study parametric change-point detection, where the goal is to identify distributional changes in time series, under local differential privacy. In the non-private setting, we derive improved finite-sample accuracy guarantees for a…

Machine Learning · Statistics 2026-02-17 Anuj Kumar Yadav , Cemre Cadir , Yanina Shkel , Michael Gastpar

In this paper, we study the problems in the discrete Fourier transform (DFT) test included in NIST SP 800-22 released by the National Institute of Standards and Technology (NIST), which is a collection of tests for evaluating both physical…

Cryptography and Security · Computer Science 2018-03-08 Hiroki Okada , Ken Umeno

The challenge of producing accurate statistics while respecting the privacy of the individuals in a sample is an important area of research. We study minimax lower bounds for classes of differentially private estimators. In particular, we…

Machine Learning · Computer Science 2024-09-19 Clément Lalanne , Aurélien Garivier , Rémi Gribonval

For some variants of regression models, including partial, measurement error or error-in-variables, latent effects, semi-parametric and otherwise corrupted linear models, the classical parametric tests generally do not perform well. Various…

Statistics Theory · Mathematics 2015-03-25 Pranab K. Sen , Jana Jureckova , Jan Picek

Tuning the hyperparameters of differentially private (DP) machine learning (ML) algorithms often requires use of sensitive data and this may leak private information via hyperparameter values. Recently, Papernot and Steinke (2022) proposed…

Machine Learning · Computer Science 2024-02-14 Antti Koskela , Tejas Kulkarni

We propose a class of kernel-based two-sample tests, which aim to determine whether two sets of samples are drawn from the same distribution. Our tests are constructed from kernels parameterized by deep neural nets, trained to maximize test…

Machine Learning · Statistics 2021-01-15 Feng Liu , Wenkai Xu , Jie Lu , Guangquan Zhang , Arthur Gretton , Danica J. Sutherland

We consider the problem of independence testing for two univariate random variables in a sequential setting. By leveraging recent developments on safe, anytime-valid inference, we propose a test with time-uniform type I error control and…

Methodology · Statistics 2024-01-29 Alexander Henzi , Michael Law

We develop a new rank-based approach for univariate two-sample testing in the presence of missing data which makes no assumptions about the missingness mechanism. This approach is a theoretical extension of the Wilcoxon-Mann-Whitney test…

Methodology · Statistics 2024-03-25 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

Due to the superior performance, large-scale pre-trained language models (PLMs) have been widely adopted in many aspects of human society. However, we still lack effective tools to understand the potential bias embedded in the black-box…

Computation and Language · Computer Science 2022-04-18 Apoorv Garg , Deval Srivastava , Zhiyang Xu , Lifu Huang

Split learning is a distributed training framework that allows multiple parties to jointly train a machine learning model over vertically partitioned data (partitioned by attributes). The idea is that only intermediate computation results,…

Machine Learning · Computer Science 2022-03-07 Xin Yang , Jiankai Sun , Yuanshun Yao , Junyuan Xie , Chong Wang