English
Related papers

Related papers: Decoding from Pooled Data: Sharp Information-Theor…

200 papers

We consider versions of the FIND algorithm where the pivot element used is the median of a subset chosen uniformly at random from the data. For the median selection we assume that subsamples of size asymptotic to $c \cdot n^\alpha$ are…

Probability · Mathematics 2013-11-20 Henning Sulzbach , Ralph Neininger , Michael Drmota

The problem of performing similarity queries on compressed data is considered. We focus on the quadratic similarity measure, and study the fundamental tradeoff between compression rate, sequence length, and reliability of queries performed…

Information Theory · Computer Science 2016-11-18 Amir Ingber , Thomas Courtade , Tsachy Weissman

In a previous paper (C. Cafaro et al., 2012), we compared an uncorrelated 3D Gaussian statistical model to an uncorrelated 2D Gaussian statistical model obtained from the former model by introducing a constraint that resembles the quantum…

Chaotic Dynamics · Physics 2015-06-17 Adom Giffin , S. A. Ali , Carlo Cafaro

Statistical Inference is the process of determining a probability distribution over the space of parameters of a model given a data set. As more data becomes available this probability distribution becomes updated via the application of…

Disordered Systems and Neural Networks · Physics 2022-04-28 David S. Berman , Jonathan J. Heckman , Marc Klinger

We consider the use of Bayesian information criteria for selection of the graph underlying an Ising model. In an Ising model, the full conditional distributions of each variable form logistic regression models, and variable selection…

Statistics Theory · Mathematics 2015-03-09 Rina Foygel Barber , Mathias Drton

The problem of individualization is recognized as crucial in almost every field. Identifying causes of effects in specific events is likewise essential for accurate decision making. However, such estimates invoke counterfactual…

Methodology · Statistics 2021-05-04 Scott Mueller , Ang Li , Judea Pearl

We study person-level differentially private (DP) mean estimation in the case where each person holds multiple samples. DP here requires the usual notion of distributional stability when $\textit{all}$ of a person's datapoints can be…

Data Structures and Algorithms · Computer Science 2024-07-22 Sushant Agarwal , Gautam Kamath , Mahbod Majid , Argyris Mouzakis , Rose Silver , Jonathan Ullman

We introduce the (private) entropy of a directed graph (in a new network coding sense) as well as a number of related concepts. We show that the entropy of a directed graph is identical to its guessing number and can be bounded from below…

Combinatorics · Mathematics 2007-11-28 Soren Riis

Data scientists have relied on samples to analyze populations of interest for decades. Recently, with the increase in the number of public data repositories, sample data has become easier to access. It has not, however, become easier to…

Databases · Computer Science 2020-01-14 Laurel Orr , Samuel Ainsworth , Walter Cai , Kevin Jamieson , Magda Balazinska , Dan Suciu

Given a network $G=(V,E)$, where each node $v$ is associated with a vector $\boldsymbol{p}_v \in \mathbb{R}^d$ representing its opinion about $d$ different topics, how can we uncover subsets of nodes that not only exhibit exceptionally high…

Social and Information Networks · Computer Science 2024-12-17 Tianyi Chen , Atsushi Miyauchi , Charalampos E. Tsourakakis

We provide novel bounds on average treatment effects (on the treated) that are valid under an unconfoundedness assumption. Our bounds are designed to be robust in challenging situations, for example, when the conditioning variables take on…

Econometrics · Economics 2026-05-12 Sokbae Lee , Martin Weidner

Suppose that we are given an arbitrary graph $G=(V, E)$ and know that each edge in $E$ is going to be realized independently with some probability $p$. The goal in the stochastic matching problem is to pick a sparse subgraph $Q$ of $G$ such…

Data Structures and Algorithms · Computer Science 2020-02-28 Soheil Behnezhad , Mahsa Derakhshan , MohammadTaghi Hajiaghayi

We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…

Machine Learning · Computer Science 2024-12-03 Maryam Aliakbarpour , Piotr Indyk , Ronitt Rubinfeld , Sandeep Silwal

Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of…

Machine Learning · Statistics 2021-03-02 Maria d'Errico , Elena Facco , Alessandro Laio , Alex Rodriguez

We propose a novel method for clustering data which is grounded in information-theoretic principles and requires no parametric assumptions. Previous attempts to use information theory to define clusters in an assumption-free way are based…

Machine Learning · Computer Science 2014-02-07 Greg Ver Steeg , Aram Galstyan , Fei Sha , Simon DeDeo

This work develops problem statements related to encoders and autoencoders with the goal of elucidating variational formulations and establishing clear connections to information-theoretic concepts. Specifically, four problems with varying…

Information Theory · Computer Science 2021-07-15 Karthik Duraisamy

We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…

Statistics Theory · Mathematics 2018-05-22 Takumi Saegusa

We describe an efficient, scalable Gaussian boson sampler based on a classical description of squeezed quantum light and a deterministic model of single-photon detectors that click when the incident amplitude falls above a given threshold.…

Quantum Physics · Physics 2025-07-24 Sarvesh Raghuraman , Aditya Patwardhan , Brian La Cour

We consider the problem of succinctly encoding a static map to support approximate queries. We derive upper and lower bounds on the space requirements in terms of the error rate and the entropy of the distribution of values over keys: our…

Data Structures and Algorithms · Computer Science 2007-10-18 David Talbot , John Talbot

The multivariate hypergeometric distribution describes sampling without replacement from a discrete population of elements divided into multiple categories. Addressing a gap in the literature, we tackle the challenge of estimating discrete…

Machine Learning · Computer Science 2024-06-11 Liam Hodgson , Danilo Bzdok