English
Related papers

Related papers: Average Bias and Polynomial Sources

200 papers

Generalized probability distributions for Maxwell-Boltzmann, Bose-Einstein and Fermi-Dirac statistics, with unequal source probabilities $q_i$ for each level $i$, are obtained by combinatorial reasoning. For equiprobable degenerate…

Statistical Mechanics · Physics 2008-08-18 Robert K. Niven , Marian Grendar

Suppose that, for any (k \geq 1), (\epsilon > 0) and sufficiently large $\sigma$, we are given a black box that allows us to sample characters from a $k$th-order Markov source over the alphabet (\{0, ..., \sigma - 1\}). Even if we know the…

Information Theory · Computer Science 2009-12-31 Travis Gagie

We introduce overdispersed black-box variational inference, a method to reduce the variance of the Monte Carlo estimator of the gradient in black-box variational inference. Instead of taking samples from the variational distribution, we use…

Machine Learning · Statistics 2016-03-04 Francisco J. R. Ruiz , Michalis K. Titsias , David M. Blei

Randomness in scientific estimation is generally assumed to arise from unmeasured or uncontrolled factors. However, when combining subjective probability estimates, heterogeneity stemming from people's cognitive or information diversity is…

Methodology · Statistics 2015-05-28 Ville Satopää , Robin Pemantle , Lyle Ungar

Percentiles and more generally, quantiles are commonly used in various contexts to summarize data. For most distributions, there is exactly one quantile that is unbiased. For distributions like the Gaussian that have the same mean and…

Methodology · Statistics 2022-01-11 Rohit Pandey

The Principle of Maximum Entropy is a rigorous technique for estimating an unknown distribution given partial information while simultaneously minimizing bias. However, an important requirement for applying the principle is that the…

Information Theory · Computer Science 2026-02-03 Kenneth Bogert , Matthew Kothe

This paper provides a framework for estimating the mean and variance of a high-dimensional normal density. The main setting considered is a fixed number of vector following a high-dimensional normal distribution with unknown mean and…

Methodology · Statistics 2019-05-07 Shyamalendu Sinha , Jeffrey D. Hart

Bayesian deep learning counts on the quality of posterior distribution estimation. However, the posterior of deep neural networks is highly multi-modal in nature, with local modes exhibiting varying generalization performance. Given a…

Machine Learning · Computer Science 2024-03-27 Bolian Li , Ruqi Zhang

Lower bounds for the average probability of error of estimating a hidden variable X given an observation of a correlated random variable Y, and Fano's inequality in particular, play a central role in information theory. In this paper, we…

Information Theory · Computer Science 2013-10-08 Flavio du Pin Calmon , Mayank Varia , Muriel Médard , Mark M. Christiansen , Ken R. Duffy , Stefano Tessaro

Scatterplots can encode a third dimension by using additional channels like size or color (e.g. bubble charts). We explore a potential misinterpretation of trivariate scatterplots, which we call the weighted average illusion, where…

Human-Computer Interaction · Computer Science 2021-08-10 Matt-Heun Hong , Jessica K. Witt , Danielle Albers Szafir

Archetypal analysis is an unsupervised learning method that uses a convex polytope to summarize multivariate data. For fixed $k$, the method finds a convex polytope with $k$ vertices, called archetype points, such that the polytope is…

Statistics Theory · Mathematics 2022-04-19 Braxton Osting , Dong Wang , Yiming Xu , Dominique Zosso

Analysis of low-degree polynomial algorithms is a powerful, newly-popular method for predicting computational thresholds in hypothesis testing problems. One limitation of current techniques for this analysis is their restriction to…

Statistics Theory · Mathematics 2020-11-10 Dmitriy Kunisky

Minimizing the relative inertia of a statistical group with respect to the inertia of the overall sample defines an unique point, the in-focus, which constitutes a context-dependent measure of typical group tendency, biased in comparison to…

Applications · Statistics 2010-04-06 François Bavaud

We continue the study of constructing explicit extractors for independent general weak random sources. The ultimate goal is to give a construction that matches what is given by the probabilistic method --- an extractor for two independent…

Computational Complexity · Computer Science 2015-03-10 Xin Li

In this paper we have proposed an almost unbiased estimator using known value of some population parameter(s) with known population proportion of an auxiliary variable. A class of estimators is defined which includes [1], [2] and [3]…

Applications · Statistics 2014-06-04 Sachin Malik , Rajesh Singh , SB Gupta

Let $f$ be analytic on $[0,1]$ with $|f^{(k)}(1/2)|\leq A\alpha^kk!$ for some constant $A$ and $\alpha<2$. We show that the median estimate of $\mu=\int_0^1f(x)\,\mathrm{d}x$ under random linear scrambling with $n=2^m$ points converges at…

Computation · Statistics 2022-07-11 Zexin Pan , Art B. Owen

Accurately measuring discrimination is crucial to faithfully assessing fairness of trained machine learning (ML) models. Any bias in measuring discrimination leads to either amplification or underestimation of the existing disparity.…

Machine Learning · Computer Science 2025-03-25 Sami Zhioua , Ruta Binkyte , Ayoub Ouni , Farah Barika Ktata

The focal-loss has become a widely used alternative to cross-entropy in class-imbalanced classification problems, particularly in computer vision. Despite its empirical success, a systematic information-theoretic study of the focal-loss…

Information Theory · Computer Science 2026-03-04 Jaimin Shah , Martina Cardone , Alex Dytso

We derive a lower bound on the smallest output entropy that can be achieved via vector quantization of a $d$-dimensional source with given expected $r$th-power distortion. Specialized to the one-dimensional case, and in the limit of…

Information Theory · Computer Science 2017-03-27 Tobias Koch , Gonzalo Vazquez-Vilar

Generalization in generative modeling is defined as the ability to learn an underlying distribution from a finite dataset and produce novel samples, with evaluation largely driven by held-out performance and perceived sample quality. In…

Machine Learning · Computer Science 2026-03-05 Jerome Garnier-Brun , Luca Biggio , Davide Beltrame , Marc Mézard , Luca Saglietti