English
Related papers

Related papers: A note on bounded distance-based information loss …

200 papers

Software defects rediscovered by a large number of customers affect various stakeholders and may: 1) hint at gaps in a software manufacturer's Quality Assurance (QA) processes, 2) lead to an over-load of a software manufacturer's support…

Software Engineering · Computer Science 2011-07-21 Andriy V. Miranskyy , Matthew Davison , Mark Reesor

We study a setting where a data holder wishes to share data with a receiver, without revealing certain summary statistics of the data distribution (e.g., mean, standard deviation). It achieves this by passing the data through a…

Cryptography and Security · Computer Science 2023-10-31 Zinan Lin , Shuaiqi Wang , Vyas Sekar , Giulia Fanti

We present a new uncertainty principle for risk-aware statistical estimation, effectively quantifying the inherent trade-off between mean squared error ($\mse$) and risk, the latter measured by the associated average predictive squared…

Information Theory · Computer Science 2021-12-13 Nikolas P. Koumpis , Dionysios S. Kalogerias

As one of the central tasks in machine learning, regression finds lots of applications in different fields. An existing common practice for solving regression problems is the mean square error (MSE) minimization approach or its regularized…

Machine Learning · Statistics 2022-11-24 Jirong Yi , Qiaosheng Zhang , Zhen Chen , Qiao Liu , Wei Shao , Yusen He , Yaohua Wang

This study proposes median consensus embedding (MCE) to address variability in low-dimensional embeddings caused by random initialization in nonlinear dimensionality reduction techniques such as $t$-distributed stochastic neighbor…

Machine Learning · Statistics 2025-12-10 Yui Tomo , Daisuke Yoneoka

Mutual Information (MI) is a crucial measure for capturing dependencies between variables, but exact computation is challenging in high dimensions with intractable likelihoods, impacting accuracy and robustness. One idea is to use an…

Machine Learning · Statistics 2025-03-13 Forough Fazeliasl , Michael Minyi Zhang , Bei Jiang , Linglong Kong

Metric differential privacy (mDP) strengthens local differential privacy (LDP) by scaling noise to semantic distance, but many machine learning (ML) systems are consumed under joint observation, where model-agnostic, per-record guarantees…

Machine Learning · Computer Science 2026-05-05 Gaoyi Chen , Minghao Li , Weishi Shi , Yan Huang , Yusheng Wei , Sourabh Yadav , Chenxi Qiu

Mutual Information (MI) is an useful tool for the recognition of mutual dependence berween data sets. Differen methods for the estimation of MI have been developed when both data sets are discrete or when both data sets are continuous. The…

Applications · Statistics 2017-08-30 Miguel A. Ré , Guillermo G. Aguirre Varela

The failure of deep neural networks to generalize to out-of-distribution data is a well-known problem and raises concerns about the deployment of trained networks in safety-critical domains such as healthcare, finance and autonomous…

Machine Learning · Computer Science 2022-06-28 Mohammed Adnan , Yani Ioannou , Chuan-Yung Tsai , Angus Galloway , H. R. Tizhoosh , Graham W. Taylor

Machine learning models are increasingly used in practice. However, many machine learning methods are sensitive to test or operational data that is dissimilar to training data. Out-of-distribution (OOD) data is known to increase the…

Machine Learning · Computer Science 2023-03-01 Tyler Cody , Laura Freeman

In machine learning, the performance of a classifier depends on both the classifier model and the separability/complexity of datasets. To quantitatively measure the separability of datasets, we create an intrinsic measure -- the…

Machine Learning · Computer Science 2021-09-14 Shuyue Guan , Murray Loew

The distance standard deviation, which arises in distance correlation analysis of multivariate data, is studied as a measure of spread. The asymptotic distribution of the empirical distance standard deviation is derived under the assumption…

Statistics Theory · Mathematics 2019-12-12 Dominic Edelmann , Donald Richards , Daniel Vogel

When the data are stored in a distributed manner, direct application of traditional statistical inference procedures is often prohibitive due to communication cost and privacy concerns. This paper develops and investigates two…

Machine Learning · Statistics 2021-08-04 Jianqing Fan , Yongyi Guo , Kaizheng Wang

Since the cost of installing and maintaining sensors is usually high, sensor locations are always strategically selected. For those aiming at inferring certain quantities of interest (QoI), it is desirable to explore the dependency between…

Computation · Statistics 2019-06-05 Xiao Lin , Asif Chowdhury , Xiaofan Wang , Gabriel Terejanu

Background: Under competing risks, the commonly used sub-distribution hazard ratio (SHR) is not easy to interpret clinically and is valid only under the proportional sub-distribution hazard (SDH) assumption. This paper introduces an…

Methodology · Statistics 2020-10-06 Jingjing Lyu , Yawen Hou , Zheng Chen

Mutual information (MI) is a fundamental measure of statistical dependence between two variables, yet accurate estimation from finite data remains notoriously difficult. No estimator is universally reliable, and common approaches fail in…

Data Analysis, Statistics and Probability · Physics 2025-10-02 Eslam Abdelaleem , K. Michael Martini , Ilya Nemenman

Statistical disclosure limitation (SDL) methods aim to provide analysts general access to a data set while limiting the risk of disclosure of individual records. Many methods in the existing literature are aimed only at the case of…

Methodology · Statistics 2015-11-03 Norman Matloff , Patrick Tendick

We propose a novel measure to assess the presence of meso-scale structures in complex networks. This measure is based on the identification of regular patterns in the adjacency matrix of the network, and on the calculation of the quantity…

Physics and Society · Physics 2015-06-18 Massimiliano Zanin , Pedro A. Sousa , Ernestina Menasalvas

Hypothesis testing is a statistical inference framework for determining the true distribution among a set of possible distributions for a given dataset. Privacy restrictions may require the curator of the data or the respondents themselves…

Information Theory · Computer Science 2017-04-28 Jiachun Liao , Lalitha Sankar , Vincent Y. F. Tan , Flavio P. Calmon

Independence analysis is an indispensable step before regression analysis to find out essential factors that influence the objects. With many applications in machine Learning, medical Learning and a variety of disciplines, statistical…

Methodology · Statistics 2022-07-08 Wenliang Pan , Yujue Li , Jianwu Liu , Pei Dang , Weixiong Mai