中文
相关论文

相关论文: Differentially private scale testing via rank tran…

200 篇论文

Testing mutual independence among multiple random variables is a fundamental problem in statistics, with wide applications in genomics, finance, and neuroscience. In this paper, we propose a new class of tests for high-dimensional mutual…

应用统计 · 统计学 2026-01-28 Ping Zhao , Huifang Ma

Background: Synthetic data has been proposed as a solution for sharing anonymized versions of sensitive biomedical datasets. Ideally, synthetic data should preserve the structure and statistical properties of the original data, while…

机器学习 · 计算机科学 2024-10-24 Ileana Montoya Perez , Parisa Movahedi , Valtteri Nieminen , Antti Airola , Tapio Pahikkala

We present the $U$-Statistic Permutation (USP) test of independence in the context of discrete data displayed in a contingency table. Either Pearson's chi-squared test of independence, or the $G$-test, are typically used for this task, but…

统计方法学 · 统计学 2022-01-19 Thomas B. Berrett , Richard J. Samworth

In this paper, we propose a general framework for distribution-free nonparametric testing in multi-dimensions, based on a notion of multivariate ranks defined using the theory of measure transportation. Unlike other existing proposals in…

统计理论 · 数学 2019-10-08 Nabarun Deb , Bodhisattva Sen

Learning effective data representations has been crucial in non-parametric two-sample testing. Common approaches will first split data into training and test sets and then learn data representations purely on the training set. However,…

机器学习 · 计算机科学 2025-05-09 Xunye Tian , Liuhua Peng , Zhijian Zhou , Mingming Gong , Arthur Gretton , Feng Liu

Ranking institutions such as medical centers or universities is based on an indicator accompanied with an uncertainty measure such as a standard deviation, and confidence intervals should be calculated to assess the quality of these ranks.…

统计方法学 · 统计学 2017-08-10 Diaa Al Mohamad , Erik W. van Zwet , Jelle J. Goeman , Aldo Solari

What proportion of treated units actually benefited from an experimental intervention? What is the median or the largest individual treatment effect? This paper develops methods for answering such questions about the distribution of…

统计方法学 · 统计学 2026-05-11 David Kim , Yongchang Su , Jake Bowers , Xinran Li

This work addresses testing the independence of two continuous and finite-dimensional random variables from the design of a data-driven partition. The empirical log-likelihood statistic is adopted to approximate the sufficient statistics of…

机器学习 · 统计学 2022-01-19 Mauricio E. Gonzalez , Jorge F. Silva , Miguel Videla , Marcos E. Orchard

Differential privacy has emerged as an significant cornerstone in the realm of scientific hypothesis testing utilizing confidential data. In reporting scientific discoveries, Bayesian tests are widely adopted since they effectively…

机器学习 · 统计学 2025-12-22 Abhisek Chakraborty , Saptati Datta

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

应用统计 · 统计学 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

We study parametric change-point detection, where the goal is to identify distributional changes in time series, under local differential privacy. In the non-private setting, we derive improved finite-sample accuracy guarantees for a…

机器学习 · 统计学 2026-02-17 Anuj Kumar Yadav , Cemre Cadir , Yanina Shkel , Michael Gastpar

In this paper, we study the problems in the discrete Fourier transform (DFT) test included in NIST SP 800-22 released by the National Institute of Standards and Technology (NIST), which is a collection of tests for evaluating both physical…

密码学与安全 · 计算机科学 2018-03-08 Hiroki Okada , Ken Umeno

The challenge of producing accurate statistics while respecting the privacy of the individuals in a sample is an important area of research. We study minimax lower bounds for classes of differentially private estimators. In particular, we…

机器学习 · 计算机科学 2024-09-19 Clément Lalanne , Aurélien Garivier , Rémi Gribonval

For some variants of regression models, including partial, measurement error or error-in-variables, latent effects, semi-parametric and otherwise corrupted linear models, the classical parametric tests generally do not perform well. Various…

统计理论 · 数学 2015-03-25 Pranab K. Sen , Jana Jureckova , Jan Picek

Tuning the hyperparameters of differentially private (DP) machine learning (ML) algorithms often requires use of sensitive data and this may leak private information via hyperparameter values. Recently, Papernot and Steinke (2022) proposed…

机器学习 · 计算机科学 2024-02-14 Antti Koskela , Tejas Kulkarni

We propose a class of kernel-based two-sample tests, which aim to determine whether two sets of samples are drawn from the same distribution. Our tests are constructed from kernels parameterized by deep neural nets, trained to maximize test…

机器学习 · 统计学 2021-01-15 Feng Liu , Wenkai Xu , Jie Lu , Guangquan Zhang , Arthur Gretton , Danica J. Sutherland

We consider the problem of independence testing for two univariate random variables in a sequential setting. By leveraging recent developments on safe, anytime-valid inference, we propose a test with time-uniform type I error control and…

统计方法学 · 统计学 2024-01-29 Alexander Henzi , Michael Law

We develop a new rank-based approach for univariate two-sample testing in the presence of missing data which makes no assumptions about the missingness mechanism. This approach is a theoretical extension of the Wilcoxon-Mann-Whitney test…

统计方法学 · 统计学 2024-03-25 Yijin Zeng , Niall M. Adams , Dean A. Bodenham

Due to the superior performance, large-scale pre-trained language models (PLMs) have been widely adopted in many aspects of human society. However, we still lack effective tools to understand the potential bias embedded in the black-box…

计算与语言 · 计算机科学 2022-04-18 Apoorv Garg , Deval Srivastava , Zhiyang Xu , Lifu Huang

Split learning is a distributed training framework that allows multiple parties to jointly train a machine learning model over vertically partitioned data (partitioned by attributes). The idea is that only intermediate computation results,…

机器学习 · 计算机科学 2022-03-07 Xin Yang , Jiankai Sun , Yuanshun Yao , Junyuan Xie , Chong Wang