English
Related papers

Related papers: Finding Non-Redundant Simpson's Paradox from Multi…

200 papers

We define a measure of redundant information based on projections in the space of probability distributions. Redundant information between random variables is information that is shared between those variables. But in contrast to mutual…

Information Theory · Computer Science 2013-05-30 Malte Harder , Christoph Salge , Daniel Polani

In search and recommendation, diversifying the multi-aspect search results could help with reducing redundancy, and promoting results that might not be shown otherwise. Many previous methods have been proposed for this task. However,…

Information Retrieval · Computer Science 2021-05-24 Jianghong Zhou , Eugene Agichtein , Surya Kallumadi

In this paper, we present a novel framework for data redundancy measurement based on probabilistic modeling of datasets, and a new criterion for redundancy detection that is resilient to noise. We also develop new methods for data…

Machine Learning · Computer Science 2024-01-17 Chunxu Cao , Qiang Zhang

Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and…

Databases · Computer Science 2020-11-20 Ruoyu Wang , Xiaobo Hu , Daniel Sun , Guoqiang Li , Raymond Wong , Shiping Chen , Jianquan Liu

Describing statistical dependencies is foundational to empirical scientific research. For uncovering intricate and possibly non-linear dependencies between a single target variable and several source variables within a system, a principled…

Information Theory · Computer Science 2024-03-28 David A. Ehrlich , Kyle Schick-Poland , Abdullah Makkeh , Felix Lanfermann , Patricia Wollstadt , Michael Wibral

Predictive multiplicity refers to the phenomenon in which classification tasks may admit multiple competing models that achieve almost-equally-optimal performance, yet generate conflicting outputs for individual samples. This presents…

Machine Learning · Computer Science 2024-02-02 Hsiang Hsu , Guihong Li , Shaohan Hu , Chun-Fu , Chen

This paper considers the problem of guessing the realization of a finite alphabet source when some side information is provided. The only knowledge the guesser has about the source and the correlated side information is that the joint…

Information Theory · Computer Science 2007-07-16 Rajesh Sundaresan

Given two input graphs, finding the largest subgraph that occurs in both, i.e., finding the maximum common subgraph, is a fundamental operator for evaluating the similarity between two graphs in graph data analysis. Existing works for…

Databases · Computer Science 2025-02-18 Kaiqiang Yu , Kaixin Wang , Cheng Long , Laks Lakshmanan , Reynold Cheng

Double descent is a surprising phenomenon in machine learning, in which as the number of model parameters grows relative to the number of data, test error drops as models grow ever larger into the highly overparameterized (data…

Complex networks often exhibit emergent behaviors, where simple dyadic interactions yield collective dynamics that cannot be explained by examining the system's units individually or in pairs. Understanding how redundant and synergistic…

Conditional-independence-based discovery uses statistical tests to identify a graphical model that represents the independence structure of variables in a dataset. These tests, however, can be unreliable, and algorithms are sensitive to…

Machine Learning · Computer Science 2026-04-21 Philipp M. Faller , Dominik Janzing

Much of social network analysis is - implicitly or explicitly - predicated on the assumption that individuals tend to be more similar to their friends than to strangers. Thus, an observed social network provides a noisy signal about the…

Social and Information Networks · Computer Science 2014-08-18 Ittai Abraham , Shiri Chechik , David Kempe , Aleksandrs Slivkins

This article precisely defines huge proofs within the system of Natural Deduction for the Minimal implicational propositional logic \mil. This is what we call an unlimited family of super-polynomial proofs. We consider huge families of…

Logic in Computer Science · Computer Science 2021-03-25 Edward Hermann Haeusler

Assume that a finite set of points is randomly sampled from a subspace of a metric space. Recent advances in computational topology have provided several approaches to recovering the geometric and topological properties of the underlying…

Algebraic Topology · Mathematics 2021-01-29 Peter Bubenik , Peter T. Kim

We study universal compression of sequences generated by monotonic distributions. We show that for a monotonic distribution over an alphabet of size $k$, each probability parameter costs essentially $0.5 \log (n/k^3)$ bits, where $n$ is the…

Information Theory · Computer Science 2007-07-13 Gil I. Shamir

The friendship paradox is a sociological phenomenon stating that most people have fewer friends than their friends do. The generalized friendship paradox refers to the same observation for attributes other than degree, and it has been…

Social and Information Networks · Computer Science 2014-11-04 Naghmeh Momeni , Michael G. Rabbat

Sequential pattern discovery is a well-studied field in data mining. Episodes are sequential patterns describing events that often occur in the vicinity of each other. Episodes can impose restrictions to the order of the events, which makes…

Databases · Computer Science 2019-04-19 Nikolaj Tatti , Boris Cule

Similarity functions measure how comparable pairs of elements are, and play a key role in a wide variety of applications, e.g., notions of Individual Fairness abiding by the seminal paradigm of Dwork et al., as well as Clustering problems.…

Machine Learning · Computer Science 2023-10-24 Leonidas Tsepenekas , Ivan Brugere , Freddy Lecue , Daniele Magazzeni

The "friendship paradox" (Feld1991) refers to the fact that, on average, people have strictly fewer friends than their friends have. I show that this over-sampling of the most popular people amplifies behaviors that involve…

Physics and Society · Physics 2017-11-21 Matthew O. Jackson

Knowledge of the association information between the attributes in a data set provides insight into the underlying structure of the data and explains the relationships (independence, synergy, redundancy) between the attributes and class (if…

Databases · Computer Science 2012-08-21 Pritam Chanda , Aidong Zhang , Murali Ramanathan