English
Related papers

Related papers: Mathematical Foundations of Data Cohesion

200 papers

Large-scale data analysis poses both statistical and computational problems which need to be addressed simultaneously. A solution is often straightforward if the data are homogeneous: one can use classical ideas of subsampling and mean…

Methodology · Statistics 2014-09-10 Peter Bühlmann , Nicolai Meinshausen

Distance covariance is a popular measure of dependence between random variables. It has some robustness properties, but not all. We prove that the influence function of the usual distance covariance is bounded, but that its breakdown value…

Methodology · Statistics 2025-08-26 Sarah Leyder , Jakob Raymaekers , Peter J. Rousseeuw

In a regression analysis, suppose we suspect that there are several heterogeneous groups in the population that a sample represents. Mixture regression models have been applied to address such problems. By modeling the conditional…

Methodology · Statistics 2013-07-02 Toshiya Hoshikawa

A broad range of on-line behaviors are mediated by interfaces in which people make choices among sets of options. A rich and growing line of work in the behavioral sciences indicate that human choices follow not only from the utility of…

Data Structures and Algorithms · Computer Science 2017-05-17 Jon Kleinberg , Sendhil Mullainathan , Johan Ugander

The likelihood function plays a pivotal role in statistical inference; it is adaptable to a wide range of models and the resultant estimators are known to have good properties. However, these results hinge on correct specification of the…

Statistics Theory · Mathematics 2017-12-15 Adam Jaeger , Nicole Lazar

Co-citation measurements can reveal the extent to which a concept representing a novel combination of existing ideas evolves towards a specialty. The strength of co-citation is represented by its frequency, which accumulates over time. Of…

Digital Libraries · Computer Science 2022-02-28 Sitaram Devarakonda , James Bradley , Dmitriy Korobskiy , Tandy Warnow , George Chacko

Correlation clustering is a flexible framework for partitioning data based solely on pairwise similarity or dissimilarity information, without requiring the number of clusters as input. However, in many practical scenarios, these pairwise…

Machine Learning · Computer Science 2025-12-11 Linus Aronsson , Morteza Haghir Chehreghani

We survey a new area of parameter-free similarity distance measures useful in data-mining, pattern recognition, learning and automatic semantics extraction. Given a family of distances on a set of objects, a distance is universal up to a…

Information Retrieval · Computer Science 2007-05-23 Paul Vitanyi

In business analytics, measure values, such as sales numbers or volumes of cargo transported, are often summed along values of one or more corresponding categories, such as time or shipping container. However, not every measure should be…

Artificial Intelligence · Computer Science 2015-12-10 Hamidreza Chinaei , Mohsen Rais-Ghasem , Frank Rudzicz

This paper introduces a comprehensive framework for complex-valued probability measures and explores their novel applications in information theory and statistical analysis. We define a complex probability measure as a phase-modulated…

Information Theory · Computer Science 2026-03-16 Siang Cheng , Hejun Xu , Tianxiao Pang

Homophily is a graph property describing the tendency of edges to connect similar nodes. There are several measures used for assessing homophily but all are known to have certain drawbacks: in particular, they cannot be reliably used for…

Machine Learning · Computer Science 2024-12-16 Mikhail Mironov , Liudmila Prokhorenkova

Retrieving cohesive subgraphs in networks is a fundamental problem in social network analysis and graph data management. These subgraphs can be used for marketing strategies or recommendation systems. Despite the introduction of numerous…

Social and Information Networks · Computer Science 2025-07-16 Dahee Kim , Song Kim , Jeongseon Kim , Junghoon Kim , Kaiyu Feng , Sungsu Lim , Jungeun Kim

Symmetry is one of the most general and useful concepts in physics. A theory or a system that has a symmetry is fundamentally constrained by it. The same constraints do not apply when the symmetry is broken. The quantitative determination…

Quantum Physics · Physics 2019-01-23 Ivan Fernandez-Corbaton

An information-theoretic development is given for the problem of compound Poisson approximation, which parallels earlier treatments for Gaussian and Poisson approximation. Let $P_{S_n}$ be the distribution of a sum $S_n=\Sumn Y_i$ of…

Probability · Mathematics 2019-06-05 A. D. Barbour , Oliver Johnson , Ioannis Kontoyiannis , Mokshay Madiman

We present a new approach to study measures on ensembles of contours, polymers or other objects interacting by some sort of exclusion condition. For concreteness we develop it here for the case of Peierls contours. Unlike existing methods,…

Probability · Mathematics 2016-08-15 Roberto Fernández , Pablo A. Ferrari , Nancy L. Garcia

Density-based directed distances -- particularly known as divergences -- between probability distributions are widely used in statistics as well as in the adjacent research fields of information theory, artificial intelligence and machine…

Statistics Theory · Mathematics 2022-03-03 Michel Broniatowski , Wolfgang Stummer

Cooperation information sharing is important to theories of human learning and has potential implications for machine learning. Prior work derived conditions for achieving optimal Cooperative Inference given strong, relatively restrictive…

Machine Learning · Computer Science 2019-02-15 Pei Wang , Pushpi Paranamana , Patrick Shafto

The spectrum and coherency are useful quantities for characterizing the temporal correlations and functional relations within and between point processes. This paper begins with a review of these quantities, their interpretation and how…

Biological Physics · Physics 2007-05-23 M. R. Jarvis , P. P. Mitra

In many applications involving multi-media data, the definition of similarity between items is integral to several key tasks, e.g., nearest-neighbor retrieval, classification, and recommendation. Data in such regimes typically exhibits…

Artificial Intelligence · Computer Science 2010-09-01 Brian McFee , Gert Lanckriet

The problem of detecting changes in covariance for a single pair of features has been studied in some detail, but may be limited in importance or general applicability. In contrast, testing equality of covariance matrices of a {\it set} of…

Methodology · Statistics 2017-12-12 Yi-Hui Zhou