Related papers: Simpson's Paradox and Collapsibility
When are inferences (whether Direct-Likelihood, Bayesian, or Frequentist) obtained from partial data valid? This paper answers this question by offering a new asymptotic theory about inference with missing data that is more general than…
We introduce and investigate a family of consequence relations with the goal of capturing certain important patterns of data-driven inference. The inspiring idea for our framework is the fact that data may reject, possibly to some degree,…
R. Duncan Luce once mentioned in a conversation that he did not consider Kolmogorov's probability theory well-constructed because it treats stochastic independence as a "numerical accident," while it should be treated as a fundamental…
We introduce the notion of a reproducible algorithm in the context of learning. A reproducible learning algorithm is resilient to variations in its samples -- with high probability, it returns the exact same output when run on two samples…
Here, by introducing a version of "Unexpected hanging paradox" we try to open a new way and a new explanation for paradoxes, similar to liar paradox. Also, we will show that we have a semantic situation which no syntactical logical system…
The modeling of probability distributions, specifically generative modeling and density estimation, has become an immensely popular subject in recent years by virtue of its outstanding performance on sophisticated data such as images and…
The contradiction of micro-reversibility and macro-irreversibility is an old problem in statistical mechanics. This article argues that irreversibility is present even at the micro level because of the mechanical-electromagnetic…
Machine learning algorithms can produce biased outcome/prediction, typically, against minorities and under-represented sub-populations. Therefore, fairness is emerging as an important requirement for the large scale application of machine…
The purpose of this work is to expand and clarify the concept of the class of Gibbs random fields and give its structure the form accepted in the theory of random processes. It is possible thanks to the proposed purely probabilistic…
In this paper I conceptualise a novel approach for capturing coincidences between events that have not necessarily an observed causal relationship. Building on the Transcendental Information Cascades approach I outline a tensor theory of…
In the present paper, we investigate consequence relations that are both paraconsistent and plausible (but still monotonic). More precisely, we put the focus on pivotal consequence relations, i.e. those relations that can be defined by a…
I give a highly selective overview of the way statistical mechanics explains the microscopic origins of the time asymmetric evolution of macroscopic systems towards equilibrium and of first order phase transitions in equilibrium. These…
We present a way of introducing joint distibution function and its marginal distribution functions for non-compatible observables. Each such marginal distribution function has the property of commutativity. Models based on this approach can…
Normality, in the colloquial sense, has historically been considered an aspirational trait, synonymous with ideality. The arithmetic average and, by extension, statistics including linear regression coefficients, have often been used to…
We introduce the coverage correlation coefficient, a novel nonparametric measure of statistical association designed to quantifies the extent to which two random variables have a joint distribution concentrated on a singular subset with…
In this paper, we argue that while the concept of a set-theoretic paradox (or paradoxical set) can be relatively well-defined within a formal setting, the concept of a set-theoretic hypodox (or hypodoxical set) remains significantly less…
Hidden variable graphical models can sometimes imply constraints on the observable distribution that are more complex than simple conditional independence relations. These observable constraints can falsify assumptions of the model that…
The aim of this paper is to show that the concept of probability is best understood by dividing this concept into two different types of probability, namely physical probability and analogical probability. Loosely speaking, a physical…
An approach to amputation, the process of introducing missing values to a complete dataset, is presented. It allows to construct missingness indicators in a flexible and principled way via copulas and Bernoulli margins and to incorporate…
Manifold hypothesis states that data points in high-dimensional space actually lie in close vicinity of a manifold of much lower dimension. In many cases this hypothesis was empirically verified and used to enhance unsupervised and…