English
Related papers

Related papers: Differentially private scale testing via rank tran…

200 papers

Using the theory of Dirichlet forms we construct a large class of continuous semimartingales on an open domain $E \subset \mathbb{R}^d$, which are governed by rank-based, in addition to name-based, characteristics. Using the results of Baur…

Probability · Mathematics 2021-04-12 David Itkin , Martin Larsson

The central question studied in this paper is Renyi Differential Privacy (RDP) guarantees for general discrete local mechanisms in the shuffle privacy model. In the shuffle model, each of the $n$ clients randomizes its response using a…

Cryptography and Security · Computer Science 2021-05-12 Antonious M. Girgis , Deepesh Data , Suhas Diggavi , Ananda Theertha Suresh , Peter Kairouz

This article introduces differentially private log-location-scale (DP-LLS) regression models, which incorporate differential privacy into LLS regression through the functional mechanism. The proposed models are established by injecting…

Machine Learning · Statistics 2024-04-16 Jiewen Sheng , Xiaolei Fang

The shuffle model of Differential Privacy (DP) has gained significant attention in privacy-preserving data analysis due to its remarkable tradeoff between privacy and utility. It is characterized by adding a shuffling procedure after each…

Combinatorics · Mathematics 2024-01-10 E Chen , Yang Cao , Yifei Ge

The Shapley value has been proposed as a solution to many applications in machine learning, including for equitable valuation of data. Shapley values are computationally expensive and involve the entire dataset. The query for a point's…

Machine Learning · Computer Science 2022-06-02 Lauren Watson , Rayna Andreeva , Hao-Tsung Yang , Rik Sarkar

Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…

Machine Learning · Statistics 2026-04-21 Antoine Chatalic , Marco Letizia , Nicolas Schreuder , Lorenzo Rosasco

Combining p-values from multiple independent tests is a fundamental task in statistical inference, but presents unique challenges when the p-values are discrete. We extend a recent optimal transport-based framework for combining discrete…

Methodology · Statistics 2025-08-05 Gonzalo Contador , Zheyang Wu

Robust classification algorithms have been developed in recent years with great success. We take advantage of this development and recast the classical two-sample test problem in the framework of classification. Based on the estimates of…

Statistics Theory · Mathematics 2019-09-18 Haiyan Cai , Bryan Goggin , Qingtang Jiang

In large scale genetic association studies, a primary aim is to test for association between genetic variants and a disease outcome. The variants of interest are often rare, and appear with low frequency among subjects. In this situation,…

Methodology · Statistics 2017-12-20 Arjun Sondhi , Kenneth Martin Rice

In this paper, we are concerned with the independence test for $k$ high-dimensional sub-vectors of a normal vector, with fixed positive integer $k$. A natural high-dimensional extension of the classical sample correlation matrix, namely…

Statistics Theory · Mathematics 2014-10-21 Zhigang Bao , Jiang Hu , Guangming Pan , Wang Zhou

Tree-based methods are powerful nonparametric techniques in statistics and machine learning. However, their effectiveness, particularly in finite-sample settings, is not fully understood. Recent applications have revealed their surprising…

Statistics Theory · Mathematics 2024-10-04 Hengrui Luo , Meng Li

Private closeness testing asks to decide whether the underlying probability distributions of two sensitive datasets are identical or differ significantly in statistical distance, while guaranteeing (differential) privacy of the data. As in…

Data Structures and Algorithms · Computer Science 2023-09-14 Clément L. Canonne , Yucheng Sun

Imbalanced learning occurs in classification settings where the distribution of class-labels is highly skewed in the training data, such as when predicting rare diseases or in fraud detection. This class imbalance presents a significant…

Machine Learning · Computer Science 2024-11-11 Lucas Rosenblatt , Yuliia Lut , Eitan Turok , Marco Avella-Medina , Rachel Cummings

In this paper, we extend the recently proposed multivariate rank energy distance, based on the theory of optimal transport, for statistical testing of distributional similarity, to soft rank energy distance. Being differentiable, this in…

Machine Learning · Statistics 2021-04-20 Shoaib Bin Masud , Boyang Lyu , Shuchin Aeron

A key challenge with machine learning approaches for ranking is the gap between the performance metrics of interest and the surrogate loss functions that can be optimized with gradient-based methods. This gap arises because ranking metrics…

Machine Learning · Computer Science 2021-11-30 Robin Swezey , Aditya Grover , Bruno Charron , Stefano Ermon

Impropriety testing for complex-valued vector has been considered lately due to potential applications ranging from digital communications to complex media imaging. This paper provides new results for such tests in the asymptotic regime,…

Signal Processing · Electrical Eng. & Systems 2020-01-07 Florent Chatelain , Nicolas Le Bihan , Jonathan H. Manton

We consider supervised learning with random decision trees, where the tree construction is completely random. The method is popularly used and works well in practice despite the simplicity of the setting, but its statistical mechanism is…

Machine Learning · Computer Science 2015-02-06 Mariusz Bojarski , Anna Choromanska , Krzysztof Choromanski , Yann LeCun

The Armitage test for linear trend in proportions can be modified using the multiple marginal model approach for three regression models with arithmetic, ordinal and logarithmic dose scores simultaneously, to be powerful against a wide…

Methodology · Statistics 2020-06-29 Ludwig A. Hothorn , Frank Schaarschmidt

We propose new statistical tests, in high-dimensional settings, for testing the independence of two random vectors and their conditional independence given a third random vector. The key idea is simple, i.e., we first transform each…

Methodology · Statistics 2026-01-28 Jinyuan Chang , Yue Du , Jing He , Qiwei Yao

Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differential privacy (DP) offers rigorous privacy guarantees and has…

Machine Learning · Computer Science 2026-01-19 Lele Zheng , Xiang Wang , Tao Zhang , Yang Cao , Ke Cheng , Yulong Shen