English
Related papers

Related papers: Generalized massive optimal data compression

200 papers

Can we analyze data without decompressing it? As our data keeps growing, understanding the time complexity of problems on compressed inputs, rather than in convenient uncompressed forms, becomes more and more relevant. Suppose we are given…

Computational Complexity · Computer Science 2018-03-05 Amir Abboud , Arturs Backurs , Karl Bringmann , Marvin Künnemann

Gaussian processes (GPs) are widely used in nonparametric regression, classification and spatio-temporal modeling, motivated in part by a rich literature on theoretical properties. However, a well known drawback of GPs that limits their use…

Methodology · Statistics 2011-06-29 Anjishnu Banerjee , David Dunson , Surya Tokdar

For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary least squares estimate in linear regression, where…

Computation · Statistics 2019-06-27 HaiYing Wang , Rong Zhu , Ping Ma

For massive data stored at multiple machines, we propose a distributed subsampling procedure for the composite quantile regression. By establishing the consistency and asymptotic normality of the composite quantile regression estimator from…

Computation · Statistics 2023-01-09 Xiaohui Yuan , Shiting Zhou , Yue Wang

In distribution compression, one aims to accurately summarize a probability distribution $\mathbb{P}$ using a small number of representative points. Near-optimal thinning procedures achieve this goal by sampling $n$ points from a Markov…

Machine Learning · Statistics 2022-10-19 Abhishek Shetty , Raaz Dwivedi , Lester Mackey

We present a smooth probabilistic reformulation of $\ell_0$ regularized regression that does not require Monte Carlo sampling and allows for the computation of exact gradients, facilitating rapid convergence to local optima of the best…

Machine Learning · Computer Science 2025-09-19 Lukas Silvester Barth , Paulo von Petersenn

The paper presents a new statistical method that enables the use of systematic errors in the maximum-likelihood regression of integer-count Poisson data to a parametric model. The method is primarily aimed at the characterization of the…

Instrumentation and Methods for Astrophysics · Physics 2024-07-18 Max Bonamente , Yang Chen , Dale Zimmerman

Performance guarantees for compression in nonlinear models under non-Gaussian observations can be achieved through the use of distributional characteristics that are sensitive to the distance to normality, and which in particular return the…

Statistics Theory · Mathematics 2017-10-03 Larry Goldstein , Xiaohan Wei

Graphical data arises naturally in several modern applications, including but not limited to internet graphs, social networks, genomics and proteomics. The typically large size of graphical data argues for the importance of designing…

Information Theory · Computer Science 2021-07-20 Payam Delgosha , Venkat Anantharam

In inference problems, we often have domain knowledge which allows us to define summary statistics that capture most of the information content in a dataset. In this paper, we present a hybrid approach, where such physics-based summaries…

Cosmology and Nongalactic Astrophysics · Physics 2025-04-24 T. Lucas Makinen , Alan Heavens , Natalia Porqueres , Tom Charnock , Axel Lapel , Benjamin D. Wandelt

Applying the standard weighted mean formula, [sum_i {n_i sigma^{-2}_i}] / [sum_i {sigma^{-2}_i}], to determine the weighted mean of data, n_i, drawn from a Poisson distribution, will, on average, underestimate the true mean by ~1 for all…

Astrophysics · Physics 2009-10-31 Kenneth J. Mighell

We consider estimation of a deterministic unknown parameter vector in a linear model with non-Gaussian noise. In the Gaussian case, dimensionality reduction via a linear matched filter provides a simple low dimensional sufficient statistic…

Applications · Statistics 2013-11-05 Jakob Vovnoboy , Ami Wiesel

A common challenge in estimating parameters of probability density functions is the intractability of the normalizing constant. While in such cases maximum likelihood estimation may be implemented using numerical integration, the approach…

Machine Learning · Statistics 2019-05-21 Shiqing Yu , Mathias Drton , Ali Shojaie

We present a way to capture high-information posteriors from training sets that are sparsely sampled over the parameter space for robust simulation-based inference. In physical inference problems, we can often apply domain knowledge to…

Machine Learning · Statistics 2025-09-26 T. Lucas Makinen , Ce Sui , Benjamin D. Wandelt , Natalia Porqueres , Alan Heavens

A compression function is a map that slims down an observational set into a subset of reduced size, while preserving its informational content. In multiple applications, the condition that one new observation makes the compressed set change…

Machine Learning · Computer Science 2024-01-09 Marco C. Campi , Simone Garatti

Subsampling is one of the popular methods to balance statistical efficiency and computational efficiency in the big data era. Most approaches aim at selecting informative or representative sample points to achieve good overall information…

Methodology · Statistics 2024-07-10 Haolin Chen , Holger Dette , Jun Yu

With increasing computing capabilities of modern supercomputers, the size of the data generated from the scientific simulations is growing rapidly. As a result, application scientists need effective data summarization techniques that can…

Human-Computer Interaction · Computer Science 2019-07-30 Soumya Dutta , Ayan Biswas , James Ahrens

Beyond the linear regime of structure formation, part of cosmological information encoded in galaxy clustering becomes inaccessible to the usual power spectrum. "Sufficient statistics", A*, were introduced recently to recapture the lost,…

Cosmology and Nongalactic Astrophysics · Physics 2015-09-30 M. Wolk , J. Carron , I. Szapudi

We derive a single pass algorithm for computing the gradient and Fisher information of Vecchia's Gaussian process loglikelihood approximation, which provides a computationally efficient means for applying the Fisher scoring algorithm for…

Computation · Statistics 2019-05-22 Joseph Guinness

Cosmic shear data contains a large amount of cosmological information encapsulated in the non-Gaussian features of the weak lensing mass maps. This information can be extracted using non-Gaussian statistics. We compare the constraining…

Cosmology and Nongalactic Astrophysics · Physics 2021-01-20 Dominik Zürcher , Janis Fluri , Raphael Sgier , Tomasz Kacprzak , Alexandre Refregier