English
Related papers

Related papers: A Dimension-Independent discriminant between distr…

200 papers

Metrics for rigorously defining a distance between two events have been used to study the properties of the dataspace manifold of particle collider physics. The probability distribution of pairwise distances on this dataspace is unique with…

High Energy Physics - Phenomenology · Physics 2025-03-07 Andrew J. Larkoski

We propose a fundamental metric for measuring the distance between two distributions. This metric, referred to as the decision-focused (DF) divergence, is tailored to stochastic linear optimization problems in which the objective…

Statistics Theory · Mathematics 2026-02-04 Suhan Liu , Mo Liu

In this article a novel approach for training deep neural networks using Bayesian techniques is presented. The Bayesian methodology allows for an easy evaluation of model uncertainty and additionally is robust to overfitting. These are…

Machine Learning · Computer Science 2019-04-03 Konstantin Posch , Jürgen Pilz

The Hierarchical Mixture of Experts (HME) is a well-known tree-based model for regression and classification, based on soft probabilistic splits. In its original formulation it was trained by maximum likelihood, and is therefore prone to…

Machine Learning · Computer Science 2012-12-12 Christopher M. Bishop , Markus Svensen

Testing mutual independence for high-dimensional observations is a fundamental statistical challenge. Popular tests based on linear and simple rank correlations are known to be incapable of detecting non-linear, non-monotone relationships,…

Statistics Theory · Mathematics 2020-02-06 Mathias Drton , Fang Han , Hongjian Shi

The weighted nearest neighbors (WNN) estimator has been popularly used as a flexible and easy-to-implement nonparametric tool for mean regression estimation. The bagging technique is an elegant way to form WNN estimators with weights…

Machine Learning · Statistics 2022-07-19 Emre Demirkaya , Yingying Fan , Lan Gao , Jinchi Lv , Patrick Vossler , Jingbo Wang

Trees and the associated shortest-path tree metrics provide a powerful framework for representing hierarchical and combinatorial structures in data. Given an arbitrary metric space, its deviation from a tree metric can be quantified by…

Machine Learning · Computer Science 2025-09-26 Pierre Houedry , Nicolas Courty , Florestan Martin-Baillon , Laetitia Chapel , Titouan Vayer

Negative binomial regression is commonly employed to analyze overdispersed count data. With small to moderate sample sizes, the maximum likelihood estimator of the dispersion parameter may be subject to a significant bias, that in turn…

Methodology · Statistics 2020-11-06 Euloge Clovis Kenne Pagui , Alessandra Salvan , Nicola Sartori

Dimension reduction plays a pivotal role in analysing high-dimensional data. However, observations with missing values present serious difficulties in directly applying standard dimension reduction techniques. As a large number of dimension…

Machine Learning · Statistics 2021-09-28 Yurong Ling , Zijing Liu , Jing-Hao Xue

Parametric adversarial divergences, which are a generalization of the losses used to train generative adversarial networks (GANs), have often been described as being approximations of their nonparametric counterparts, such as the…

Machine Learning · Computer Science 2021-10-22 Gabriel Huang , Hugo Berard , Ahmed Touati , Gauthier Gidel , Pascal Vincent , Simon Lacoste-Julien

In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these…

Statistics Theory · Mathematics 2015-06-03 Rahul Agarwal , Pierre Sacre , Sridevi V. Sarma

This paper investigates the comparative performance of two fundamental approaches to solving linear regression problems: the closed-form Moore-Penrose pseudoinverse and the iterative gradient descent method. Linear regression is a…

Machine Learning · Computer Science 2025-05-30 Alex Adams

In this article, we investigate posterior convergence of nonparametric binary and Poisson regression under possible model misspecification, assuming general stochastic process prior with appropriate properties. Our model setup and objective…

Statistics Theory · Mathematics 2020-05-04 Debashis Chatterjee , Sourabh Bhattacharya

This article presents a general framework for the transport of probability measures towards minimum divergence generative modeling and sampling using ordinary differential equations (ODEs) and Reproducing Kernel Hilbert Spaces (RKHSs),…

Machine Learning · Statistics 2024-02-14 Biraj Pandey , Bamdad Hosseini , Pau Batlle , Houman Owhadi

Mean-Field is an efficient way to approximate a posterior distribution in complex graphical models and constitutes the most popular class of Bayesian variational approximation methods. In most applications, the mean field distribution…

Machine Learning · Computer Science 2015-02-23 Pierre Baqué , Jean-Hubert Hours , François Fleuret , Pascal Fua

The problem of estimating a high-dimensional sparse vector $\boldsymbol{\theta} \in \mathbb{R}^n$ from an observation in i.i.d. Gaussian noise is considered. The performance is measured using squared-error loss. An empirical Bayes shrinkage…

Information Theory · Computer Science 2018-12-31 Pavan Srinath , Ramji Venkataramanan

The statistics and machine learning communities have recently seen a growing interest in classification-based approaches to two-sample testing. The outcome of a classification-based two-sample test remains a rejection decision, which is not…

Statistics Theory · Mathematics 2022-11-15 Loris Michel , Jeffrey Näf , Nicolai Meinshausen

Binary density ratio estimation (DRE), the problem of estimating the ratio $p_1/p_2$ given their empirical samples, provides the foundation for many state-of-the-art machine learning algorithms such as contrastive representation learning…

Machine Learning · Computer Science 2021-12-08 Lantao Yu , Yujia Jin , Stefano Ermon

In this paper, we study classification and regression error bounds for inhomogenous data that are independent but not necessarily identically distributed. First, we consider classification of data in the presence of non-stationary noise and…

Information Theory · Computer Science 2024-04-04 Ghurumuruhan Ganesan

This paper considers multiple binary hypothesis tests with adaptive allocation of sensing resources from a shared budget over a small number of stages. A Bayesian formulation is provided for the multistage allocation problem of minimizing…

Methodology · Statistics 2014-11-05 Dennis Wei