Related papers: Measuring sets by means
Classification systems are evaluated in a countless number of papers. However, we find that evaluation practice is often nebulous. Frequently, metrics are selected without arguments, and blurry terminology invites misconceptions. For…
We discuss two main ways in comparing and evaluating the size of sets: the "Cantorian" way, grounded on the so called Hume principle (two sets have equal size if they are equipotent), and the "Euclidean" way, maintaining Euclid's principle…
Data sets for fairness relevant tasks can lack examples or be biased according to a specific label in a sensitive attribute. We demonstrate the usefulness of weight based meta-learning approaches in such situations. For models that can be…
We consider the problem of meta-analyzing two-group studies that report the median of the outcome. Often, these studies are excluded from meta-analysis because there are no well-established statistical methods to pool the difference of…
Organising the relevant literature and by letting statistical convergence play the main role in the theory of compactness, a variant of compactness called statistical compactness has been achieved. As in case of sequential compactness, one…
We show that in doubling, geodesic metric measure spaces (including, for example, Euclidean space), sets of positive measure have a certain large-scale metric density property. As an application, we prove that a set of positive measure in…
Cluster analysis is a fundamental research issue in statistics and machine learning. In many modern clustering methods, we need to determine whether two subsets of samples come from the same cluster. Since these subsets are usually…
Apportionment is the task of assigning resources to entities with different entitlements in a fair manner, and specifically a manner that is as proportional as possible. The best-known application is the assignment of parliamentary seats to…
We study finitely additive extensions of the asymptotic density to all the subsets of natural numbers. Such measures are called density measures. We consider a class of density measures constructed from free ultrafilters on $\mathbb{N}$ and…
Starting with a set of weighted items, we want to create a generic sample of a certain size that we can later use to estimate the total weight of arbitrary subsets. For this purpose, we propose priority sampling which tested on Internet…
When factorizing binary matrices, we often have to make a choice between using expensive combinatorial methods that retain the discrete nature of the data and using continuous methods that can be more efficient but destroy the discrete…
Many real-world classification problems are significantly class-imbalanced to detriment of the class of interest. The standard set of proper evaluation metrics is well-known but the usual assumption is that the test dataset imbalance equals…
The paper has two main goals. The first is to take a new approach to rearrangements on certain classes of measurable real-valued functions on $\mathbb{R}^n$. Rearrangements are maps that are monotonic (up to sets of measure zero) and…
Quantization for a probability distribution refers to the idea of estimating a given probability by a discrete probability supported by a finite number of points. In this paper, firstly a general approach to this process is outlined using…
In this chapter, I review the main methods and techniques of complex systems science. As a first step, I distinguish among the broad patterns which recur across complex systems, the topics complex systems science commonly studies, the tools…
For a large class of statistical systems a geometric mean value of the observables is constrained. These observables are characterized by a power-law statistical distribution.
We introduce a rigorous framework for the quantification of coherence and identify intuitive and easily computable measures of coherence. We achieve this by adopting the viewpoint of coherence as a physical resource. By determining defining…
We investigate the set of limit points of averages of rearrangements of a given sequence. We study how the properties of the sequence determine the structure of that set and what type of sets we can expect as the set of such accessible…
We study random composite structures considered up to symmetry that are sampled according to weights on the inner and outer structures. This model may be viewed as an unlabelled version of Gibbs partitions and encompasses multisets of…
Quantization for probability distributions refers broadly to estimating a given probability measure by a discrete probability measure supported by a finite number of points. We consider general geometric approaches to quantization using…