English
Related papers

Related papers: An algorithmic and a geometric characterization of…

200 papers

Learning the distribution of a continuous or categorical response variable $\boldsymbol y$ given its covariates $\boldsymbol x$ is a fundamental problem in statistics and machine learning. Deep neural network-based supervised learning…

Machine Learning · Statistics 2022-12-07 Xizewen Han , Huangjie Zheng , Mingyuan Zhou

Multivariate information theory provides a general and principled framework for understanding how the components of a complex system are connected. Existing analyses are coarse in nature -- built up from characterizations of discrete…

Information Theory · Computer Science 2025-05-30 Kieran A. Murphy , Yujing Zhang , Dani S. Bassett

We formulate the statistics of the discrete multicomponent fragmentation event using a methodology borrowed from statistical mechanics. We generate the ensemble of all feasible distributions that can be formed when a single integer…

Statistical Mechanics · Physics 2020-07-03 Themis Matsoukas

We give a principled method for decomposing the predictive uncertainty of a model into aleatoric and epistemic components with explicit semantics relating them to the real-world data distribution. While many works in the literature have…

Machine Learning · Computer Science 2024-12-30 Gustaf Ahdritz , Aravind Gollakota , Parikshit Gopalan , Charlotte Peale , Udi Wieder

Conditional autoregressive (CAR) models are commonly used to capture spatial correlation in areal unit data, and are typically specified as a prior distribution for a set of random effects, as part of a hierarchical Bayesian model. The…

Applications · Statistics 2012-05-17 Duncan Lee , Richard Mitchell

Hypergraphs are structures that can be decomposed or described; in other words they are recursively countable. Here, we get exact and asymptotic enumeration results on hypergraphs by means of exponential generating functions. The number of…

Discrete Mathematics · Computer Science 2008-06-20 Tsiriniaina Andriamampianina

We build information geometry for a partially ordered set of variables and define the orthogonal decomposition of information theoretic quantities. The natural connection between information geometry and order theory leads to efficient…

Information Theory · Computer Science 2016-11-18 Mahito Sugiyama , Hiroyuki Nakahara , Koji Tsuda

Variable selection is a difficult problem that is particularly challenging in the analysis of high-dimensional genomic data. Here, we introduce the CAR score, a novel and highly effective criterion for variable ranking in linear regression…

Methodology · Statistics 2011-07-20 Verena Zuber , Korbinian Strimmer

A coarse description of a subset A of omega is a subset D of omega such that the symmetric difference of A and D has asymptotic density 0. We study the extent to which noncomputable information can be effectively recovered from all coarse…

Logic · Mathematics 2015-05-08 Denis R. Hirschfeldt , Carl G. Jockusch , Rutger Kuyper , Paul E. Schupp

We consider covariate adjusted regression (CAR), a regression method for situations where predictors and response are observed after being distorted by a multiplicative factor. The distorting factors are unknown functions of an observable…

Statistics Theory · Mathematics 2016-08-16 Damla Şentürk , Hans-Georg Müller

We propose a data-driven, coarse-graining formulation in the context of equilibrium statistical mechanics. In contrast to existing techniques which are based on a fine-to-coarse map, we adopt the opposite strategy by prescribing a…

Machine Learning · Statistics 2017-02-01 Markus Schöberl , Nicholas Zabaras , Phaedon-Stelios Koutsourelakis

Conformal Prediction (CP) is a principled framework for quantifying uncertainty in blackbox learning models, by constructing prediction sets with finite-sample coverage guarantees. Traditional approaches rely on scalar nonconformity scores,…

Machine Learning · Statistics 2025-05-07 Gauthier Thurin , Kimia Nadjahi , Claire Boyer

Nonparametric estimation of the conditional distribution of a response given high-dimensional features is a challenging problem. It is important to allow not only the mean but also the variance and shape of the response density to change…

Machine Learning · Statistics 2013-12-05 Francesca Petralia , Joshua Vogelstein , David B. Dunson

Let $\mathfrak{C}$ be a class of probability distributions over the discrete domain $[n] = \{1,...,n\}.$ We show that if $\mathfrak{C}$ satisfies a rather general condition -- essentially, that each distribution in $\mathfrak{C}$ can be…

Machine Learning · Computer Science 2012-10-03 Siu-on Chan , Ilias Diakonikolas , Rocco A. Servedio , Xiaorui Sun

We develop a model to describe the properties of random assemblies of polydisperse hard spheres. We show that the key features to describe the system are (i) the dependence between the free volume of a sphere and the various coordination…

Disordered Systems and Neural Networks · Physics 2015-05-19 Maximilien Danisch , Yuliang Jin , Hernan A. Makse

The aim in many sciences is to understand the mechanisms that underlie the observed distribution of variables, starting from a set of initial hypotheses. Causal discovery allows us to infer mechanisms as sets of cause and effect…

Machine Learning · Computer Science 2025-03-05 Ashka Shah , Adela DePavia , Nathaniel Hudson , Ian Foster , Rick Stevens

Feature selection can facilitate the learning of mixtures of discrete random variables as they arise, e.g. in crowdsourcing tasks. Intuitively, not all workers are equally reliable but, if the less reliable ones could be eliminated, then…

Machine Learning · Statistics 2017-11-28 Vincent Zhao , Steven W. Zucker

A model of a geometric algorithm is introduced and methodology of its operation is presented for the dynamic partitioning of data spaces.

Data Structures and Algorithms · Computer Science 2014-12-30 Christopher A. Tucker

This paper gives poly-logarithmic-round, distributed D-approximation algorithms for covering problems with submodular cost and monotone covering constraints (Submodular-cost Covering). The approximation ratio D is the maximum number of…

Data Structures and Algorithms · Computer Science 2020-05-29 Christos Koufogiannakis , Neal E. Young

Recently, there has been an increasing interest in designing distributed convex optimization algorithms under the setting where the data matrix is partitioned on features. Algorithms under this setting sometimes have many advantages over…

Machine Learning · Computer Science 2016-12-05 Zihao Chen , Luo Luo , Zhihua Zhang