English
Related papers

Related papers: PKLM: A flexible MCAR test using Classification

200 papers

In this letter, we propose a novel statistical method to measure which system is better suited to probe small deviations from the usual quantum behavior. Such deviations are motivated by a number of theoretical and phenomenological…

Pre-trained machine learning (ML) predictions have been increasingly used to complement incomplete data to enable downstream scientific inquiries, but their naive integration risks biased inferences. Recently, multiple methods have been…

Methodology · Statistics 2025-11-12 Xingran Chen , Tyler McCormick , Bhramar Mukherjee , Zhenke Wu

We study the problem of distributed Kalman filtering for sensor networks in the presence of model uncertainty. More precisely, we assume that the actual state-space model belongs to a ball, in the Kullback-Leibler topology, about the…

Optimization and Control · Mathematics 2020-04-20 Mattia Zorzi

Missing data can lead to inefficiencies and biases in analyses, in particular when data are missing not at random (MNAR). It is thus vital to understand and correctly identify the missing data mechanism. Recovering missing values through a…

Methodology · Statistics 2022-12-08 Jack Noonan , Adetola Adedamola Adediran , Robin Mitra , Stefanie Biedermann

We consider the fundamental problem of estimating a discrete distribution on a domain of size $K$ with high probability in Kullback-Leibler divergence. We provide upper and lower bounds on the minimax estimation rate, which show that the…

Machine Learning · Statistics 2026-02-23 Dirk van der Hoeven , Julia Olkhovskaia , Tim van Erven

Estimating Kullback Leibler (KL) divergence from samples of two distributions is essential in many machine learning problems. Variational methods using neural network discriminator have been proposed to achieve this task in a scalable…

Machine Learning · Computer Science 2021-10-01 Sandesh Ghimire , Aria Masoomi , Jennifer Dy

Missing data are a common problem for both the construction and implementation of a prediction algorithm. Pattern mixture kernel submodels (PMKS) - a series of submodels for every missing data pattern that are fit using only data from that…

Methodology · Statistics 2017-04-27 Sarah Fletcher Mercaldo , Jeffrey D. Blume

Empirical risk minimization, a cornerstone in machine learning, is often hindered by the Optimizer's Curse stemming from discrepancies between the empirical and true data-generating distributions.To address this challenge, the robust…

Machine Learning · Computer Science 2024-08-20 Haojie Yan , Minglong Zhou , Jiayi Guo

The problem of predicting independent Poisson random variables is commonly encountered in real-life practice. Simultaneous predictive distributions for independent Poisson observables are investigated, and the performance of predictive…

Statistics Theory · Mathematics 2023-12-06 Xiao Li , Fumiyasu Komaki

Data collection is a critical step in statistical inference and data science, and the goal of statistical experimental design (ED) is to find the data collection setup that can provide most information for the inference. In this work we…

Computation · Statistics 2020-07-01 Ziqiao Ao , Jinglai Li

This paper proposes two linear projection methods for supervised dimension reduction using only the first and second-order statistics. The methods, each catering to a different parameter regime, are derived under the general Gaussian model…

Information Theory · Computer Science 2024-08-13 Biao Chen , Joshua Kortje

Estimating the Kullback-Leibler (KL) divergence between two distributions given samples from them is well-studied in machine learning and information theory. Motivated by considerations of multi-group fairness, we seek KL divergence…

Machine Learning · Computer Science 2022-03-01 Parikshit Gopalan , Nina Narodytska , Omer Reingold , Vatsal Sharan , Udi Wieder

Maximizing the Kullback-Leibler divergence (KLD) is a fundamental problem in waveform design for active sensing and hypothesis testing, as it directly relates to the error exponent of detection probability. However, the associated…

Signal Processing · Electrical Eng. & Systems 2026-01-05 Jeongwoo Park , Seongkyu Jung , Kaiming Shen , Jeonghun Park

A goodness-of-fit test for the Functional Linear Model with Scalar Response (FLMSR) with responses Missing at Random (MAR) is proposed in this paper. The test statistic relies on a marked empirical process indexed by the projected…

Shape constraints yield flexible middle grounds between fully nonparametric and fully parametric approaches to modeling distributions of data. The specific assumption of log-concavity is motivated by applications across economics, survival…

Methodology · Statistics 2024-04-16 Robin Dunn , Aditya Gangrade , Larry Wasserman , Aaditya Ramdas

Recently in [1, 2], Ali-Akbar Bromideh introduced the Kullback-Leibler Divergence (KLD) test statistic in discrim- inating between two models. It was found that the Ratio Minimized Kulback-Leibler Divergence (RMKLD) works better than the…

Methodology · Statistics 2017-10-02 Papa Ngom , Jean de Dieu Nkurunziza , Carlos Simplice Ogouyandjou

Knowing if a model will generalize to data 'in the wild' is crucial for safe deployment. To this end, we study model disagreement notions that consider the full predictive distribution - specifically disagreement based on Hellinger…

Machine Learning · Computer Science 2023-12-14 Mona Schirmer , Dan Zhang , Eric Nalisnick

Maximum Likelihood Estimators (MLE) has many good properties. For example, the asymptotic variance of MLE solution attains equality of the asymptotic Cram{\'e}r-Rao lower bound (efficiency bound), which is the minimum possible variance for…

Machine Learning · Statistics 2019-11-05 Song Liu , Takafumi Kanamori , Wittawat Jitkrittum , Yu Chen

The goal of positive-unlabeled (PU) learning is to train a binary classifier on the basis of training data containing positive and unlabeled instances, where unlabeled observations can belong either to the positive class or to the negative…

Machine Learning · Statistics 2024-04-02 Paweł Teisseyre , Konrad Furmańczyk , Jan Mielniczuk

Relational models generalize log-linear models to arbitrary discrete sample spaces by specifying effects associated with any subsets of their cells. A relational model may include an overall effect, pertaining to every cell after a…

Methodology · Statistics 2019-05-17 Anna Klimova , Tamás Rudas