English
Related papers

Related papers: Maximum Entropy Based Significance of Itemsets

200 papers

Given a knowledge base KB containing first-order and statistical facts, we consider a principled method, called the random-worlds method, for computing a degree of belief that some formula Phi holds given KB. If we are reasoning about a…

Artificial Intelligence · Computer Science 2009-09-25 A. J. Grove , J. Y. Halpern , D. Koller

Here, we investigate the uncertainty of dynamical observables in classical systems manipulated by repeated measurements and feedback control; the precision should be enhanced in the presence of an external controller but limited by the…

Statistical Mechanics · Physics 2020-01-23 Tan Van Vu , Yoshihiko Hasegawa

We seek an entropy estimator for discrete distributions with fully empirical accuracy bounds. As stated, this goal is infeasible without some prior assumptions on the distribution. We discover that a certain information moment assumption…

Information Theory · Computer Science 2022-12-27 Doron Cohen , Aryeh Kontorovich , Aaron Koolyk , Geoffrey Wolfer

Maximum entropy (Maxent) models are a class of statistical models that use the maximum entropy principle to estimate probability distributions from data. Due to the size of modern data sets, Maxent models need efficient optimization…

Machine Learning · Statistics 2024-03-12 Gabriel P. Langlois , Jatan Buch , Jérôme Darbon

Entropy is a measure of self-information which is used to quantify losses. Entropy was developed in thermodynamics, but is also used to compare probabilities based on their deviating information content. Corresponding model uncertainty is…

Probability · Mathematics 2018-01-23 Alois Pichler , Ruben Schlotter

In a variety of applications it is important to extract information from a probability measure $\mu$ on an infinite dimensional space. Examples include the Bayesian approach to inverse problems and possibly conditioned) continuous time…

Probability · Mathematics 2016-06-02 Frank Pinski , Gideon Simpson , Andrew Stuart , Hendrik Weber

We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise senses, including a kind of instance-optimality. Our data-dependent convergence guarantees…

Statistics Theory · Mathematics 2024-02-14 Aryeh Kontorovich , Amichai Painsky

We define the concept of dependence among multiple variables using maximum entropy techniques and introduce a graphical notation to denote the dependencies. Direct inference of information theoretic quantities from data uncovers…

Quantitative Methods · Quantitative Biology 2007-07-13 Ilya Nemenman

A finite set is "hidden" if its elements are not directly enumerable or if its size cannot be ascertained via a deterministic query. In public health, epidemiology, demography, ecology and intelligence analysis, researchers have developed a…

Statistics Theory · Mathematics 2019-10-17 Si Cheng , Daniel J. Eck , Forrest W. Crawford

In many applications in biology, engineering and economics, identifying similarities and differences between distributions of data from complex processes requires comparing finite categorical samples of discrete counts. Statistical…

Methodology · Statistics 2023-07-11 Francesco Camaglia , Ilya Nemenman , Thierry Mora , Aleksandra M. Walczak

Density-based directed distances -- particularly known as divergences -- between probability distributions are widely used in statistics as well as in the adjacent research fields of information theory, artificial intelligence and machine…

Statistics Theory · Mathematics 2022-03-03 Michel Broniatowski , Wolfgang Stummer

Mining frequent patterns is plagued by the problem of pattern explosion making pattern reduction techniques a key challenge in pattern mining. In this paper we propose a novel theoretical framework for pattern reduction. We do this by…

Databases · Computer Science 2019-04-25 Nikolaj Tatti , Fabian Moerchen , Toon Calders

The task of item recommendation requires ranking a large catalogue of items given a context. Item recommendation algorithms are evaluated using ranking metrics that depend on the positions of relevant items. To speed up the computation of…

Information Retrieval · Computer Science 2019-12-06 Steffen Rendle

In many problems in data mining and machine learning, data items that need to be clustered or classified are not points in a high-dimensional space, but are distributions (points on a high dimensional simplex). For distributions, natural…

Data Structures and Algorithms · Computer Science 2007-07-13 Sudipto Guha , Andrew McGregor , Suresh Venkatasubramanian

Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the…

Machine Learning · Statistics 2016-01-21 Simone Romano , Nguyen Xuan Vinh , James Bailey , Karin Verspoor

Feature-importance methods show promise in transforming machine learning models from predictive engines into tools for scientific discovery. However, due to data sampling and algorithmic stochasticity, expressive models can be unstable,…

Machine Learning · Statistics 2026-05-29 Joseph Paillard , Angel Reyero Lobo , Denis A. Engemann , Bertrand Thirion

Ranked set sampling is a sampling design which has a wide range of applications in industrial statistics, and environmental and ecological studies, etc.. It is well known that ranked set samples provide more Fisher information than simple…

Statistics Theory · Mathematics 2013-01-21 Mohammad Jafari Jozani , Jafar Ahmadi

Measuring dependence between two events, or equivalently between two binary random variables, amounts to expressing the dependence structure inherent in a $2\times 2$ contingency table in a real number between $-1$ and $1$. Countless such…

Methodology · Statistics 2025-11-13 Marc-Oliver Pohle , Timo Dimitriadis , Jan-Lukas Wermuth

We derive an optimal bound on the sum of entropic uncertainties of two or more observables when they are sequentially measured on the same ensemble of systems. This optimal bound is shown to be greater than or equal to the bounds derived in…

Quantum Physics · Physics 2009-11-07 M. D. Srinivas

An efficient approach to the calculation of the $\epsilon$-entropy is proposed. The method is based on the idea of looking at the information content of a string of data, by analyzing the signal only at the instants when the fluctuations…

chao-dyn · Physics 2009-10-31 M. Abel , L. Biferale , M. Cencini , M. Falcioni , D. Vergni , A. Vulpiani