English
Related papers

Related papers: Mathematical Foundations of Data Cohesion

200 papers

A method is discussed that allows combining sets of differential or inclusive measurements. It is assumed that at least one measurement was obtained with simultaneously fitting a set of nuisance parameters, representing sources of…

Data Analysis, Statistics and Probability · Physics 2018-01-09 Jan Kieseler

We study the problem of reconstructing the probability measure of the Curie-Weiss model from a sample of the voting behaviour of a subset of the population. While originally used to study phase transitions in statistical mechanics, the…

Probability · Mathematics 2025-08-06 Miguel Ballesteros , Ivan Naumkin , Gabor Toth

The Jaccard similarity index has often been employed in science and technology as a means to quantify the similarity between two sets. When modified to operate on real-valued values, the Jaccard similarity index can be applied to compare…

Data Analysis, Statistics and Probability · Physics 2024-10-23 Gonzalo Travieso , Alexandre Benatti , Luciano da F. Costa

We present a unified framework for quantifying the similarity between representations through the lens of \textit{usable} information, offering a rigorous theoretical and empirical synthesis across three key dimensions. First, addressing…

Machine Learning · Computer Science 2026-05-29 Antonio Almudévar , Alfonso Ortega

In movement ecology, the few works that have taken collective behaviour into account are data-driven and rely on simplistic theoretical assumptions, relying in metrics that may or may not be measuring what is intended. In the present paper,…

Quantitative Methods · Quantitative Biology 2019-03-18 Rocio Joo , Marie-Pierre Etienne , Nicolas Bez , Stéphanie Mahévas

Measuring atomic and molecular interactions was one of the main objectives of physics during the past century. It was an essential step not only in itself but because most macroscopic properties can be derived once one knows interaction…

Biological Physics · Physics 2014-06-16 Z. Di , M. Gho , X. Lu , G. Li , B. M. Roehner , N. J. Suematsu , C. Yepremian

The concept of entropy, firstly introduced in information theory, rapidly became popular in many applied sciences via Shannon's formula to measure the degree of heterogeneity among observations. A rather recent research field aims at…

Methodology · Statistics 2017-03-20 Linda Altieri , Daniela Cocchi , Giulia Roli

The development of science has been transforming man's view towards nature for centuries. Observing structures and patterns in an effective approach to discover regularities from data is a key step toward theory-building. With increasingly…

Computational Physics · Physics 2025-06-09 Guang-Xing Li

Aggregation-diffusion equations are foundational tools for modelling biological aggregations. Their principal use is to link the collective movement mechanisms of organisms to their emergent space use patterns in a concrete mathematical…

Populations and Evolution · Quantitative Biology 2025-04-16 Jonathan R. Potts

We propose a statistical framework to quantify location and co-location associations of economic activities using information-theoretic measures. We relate the resulting measures to existing measures of revealed comparative advantage,…

Applications · Statistics 2020-04-23 Alje van Dam , Andres Gomez-Lievano , Frank Neffke , Koen Frenken

Quantifying cooperation or synergy among random variables in predicting a single target random variable is an important problem in many complex systems. We review three prior information-theoretic measures of synergy and introduce a novel…

Information Theory · Computer Science 2014-04-02 Virgil Griffith , Christof Koch

A computational theory for clustering and a semi-supervised clustering algorithm is presented. Clustering is defined to be the obtainment of groupings of data such that each group contains no anomalies with respect to a chosen grouping…

Machine Learning · Computer Science 2025-07-17 Nassir Mohammad

In classical density (or density-functional) estimation, it is standard to assume that the underlying distribution has a density with respect to the Lebesgue measure. However, when the data distribution is a mixture of continuous and…

Methodology · Statistics 2025-08-05 Aytijhya Saha , Aaditya Ramdas

In scientific inference problems, the underlying statistical modeling assumptions have a crucial impact on the end results. There exist, however, only a few automatic means for validating these fundamental modelling assumptions. The…

Methodology · Statistics 2019-05-21 Andreas Svensson , Dave Zachariah , Petre Stoica , Thomas B. Schön

The subject of features normalization plays an important central role in data representation, characterization, visualization, analysis, comparison, classification, and modeling, as it can substantially influence and be influenced by all of…

Machine Learning · Computer Science 2024-09-18 Alexandre Benatti , Luciano da F. Costa

A key ingredient in social contagion dynamics is reinforcement, as adopting a certain social behavior requires verification of its credibility and legitimacy. Memory of non-redundant information plays an important role in reinforcement,…

Physics and Society · Physics 2015-07-29 Wei Wang , Ming Tang , Hai-Feng Zhang , Ying-Cheng Lai

Consider an unlimited homogeneous medium disturbed by points generated via Poisson process. The neighborhood of a point plays an important role in spatial statistics problems. Here, we obtain analytically the distance statistics to $k$th…

Statistical Mechanics · Physics 2015-08-11 Cristiano Roberto Fabri Granzotti , Alexandre Souto Martinez

We propose a novel measure of statistical depth, the metric spatial depth, for data residing in an arbitrary metric space. The measure assigns high (low) values for points located near (far away from) the bulk of the data distribution,…

Statistics Theory · Mathematics 2023-06-19 Joni Virta

Several data analysis techniques employ similarity relationships between data points to uncover the intrinsic dimension and geometric structure of the underlying data-generating mechanism. In this paper we work under the model assumption…

Machine Learning · Statistics 2019-04-09 Nicolas Garcia Trillos , Daniel Sanz-Alonso , Ruiyi Yang

Multivariate datasets are common in various real-world applications. Recently, copulas have received significant attention for modeling dependencies among random variables. A copula-based information measure is required to quantify the…

Methodology · Statistics 2024-08-06 Mohd. Arshad , Swaroop Georgy Zachariah , Ashok Kumar Pathak