English
Related papers

Related papers: Efficient estimation of the cardinality of large d…

200 papers

In recent years there has been a growing interest in developing "streaming algorithms" for efficient processing and querying of continuous data streams. These algorithms seek to provide accurate results while minimizing the required storage…

Data Structures and Algorithms · Computer Science 2016-06-06 Reuven Cohen , Liran Katzir , Aviv Yehezkel

In this paper, we use the Poincare separation theorem for estimating the eigenvalues of the fine grid. We propose a randomized version of the algorithm where several different coarse grids are constructed thus leading to more comprehensive…

Numerical Analysis · Mathematics 2015-03-20 Pawan Kumar

In this paper, we analyze the celebrated EM algorithm from the point of view of proximal point algorithms. More precisely, we study a new type of generalization of the EM procedure introduced in \cite{Chretien&Hero:98} and called…

Computation · Statistics 2012-01-31 Stéphane Chrétien , Alfred O. Hero

An algorith to count, or alternatively generate, all k-element transversals of a set system is presented and compared with three known methods. For special cases it works in output-linear time.

Discrete Mathematics · Computer Science 2017-03-01 Marcel Wild

Recent studies have demonstrated that correntropy is an efficient tool for analyzing higher-order statistical moments in nonGaussian noise environments. Although correntropy has been used with complex data, no theoretical study was pursued…

Information Theory · Computer Science 2016-08-19 João Paulo Ferreira Guimarães

We develop two surprising new results regarding the use of proper scoring rules for evaluating the predictive quality of two alternative sequential forecast distributions. Both of the proponents prefer to be awarded a score derived from the…

Probability · Mathematics 2019-09-17 Frank Lad , Giuseppe Sanfilippo

This study proposes a data condensation method for multivariate kernel density estimation by genetic algorithm. First, our proposed algorithm generates multiple subsamples of a given size with replacement from the original sample. The…

Methodology · Statistics 2022-03-04 Kiheiji Nishida

We examine the estimation of the Kullback-Leibler (KL) divergence and the use of the goodness-of-fit test for multivariate continuous distributions. Our starting point is the maximum entropy principle for Shannon entropy: among all…

Statistics Theory · Mathematics 2026-03-10 Mehmet Siddik Cadirci , Martin Singull

Learning interpretable models has become a major focus of machine learning research, given the increasing prominence of machine learning in socially important decision-making. Among interpretable models, rule lists are among the best-known…

Machine Learning · Computer Science 2024-06-19 Leonardo Pellegrina , Fabio Vandin

We present a simple Coulomb gas method to calculate analytically the probability of rare events where the maximum eigenvalue of a random matrix is much larger than its typical value. The large deviation function that characterizes this…

Statistical Mechanics · Physics 2009-02-27 Satya N. Majumdar , Massimo Vergassola

\emph{$K$-best enumeration}, which asks to output $k$-best solutions without duplication, is a helpful tool in data analysis for many fields. In such fields, graphs typically represent data. Thus subgraph enumeration has been paid much…

Data Structures and Algorithms · Computer Science 2024-05-14 Kazuhiro Kurita , Kunihiro Wasa

In this research paper, we address the Distinct Elements estimation problem in the context of streaming algorithms. The problem involves estimating the number of distinct elements in a given data stream $\mathcal{A} = (a_1, a_2,\ldots,…

Data Structures and Algorithms · Computer Science 2023-06-13 Mridul Nandi , Soumit Paul

According to a general probabilistic principle, the natural divisors of friable integers (i.e.~free of large prime factors) should normally present a Gaussian distribution. We show that this indeed is the case with conditional density…

Number Theory · Mathematics 2018-05-29 Sary Drappeau , Gérald Tenenbaum

Human ratings have become a crucial resource for training and evaluating machine learning systems. However, traditional elicitation methods for absolute and comparative rating suffer from issues with consistency and often do not distinguish…

Human-Computer Interaction · Computer Science 2021-08-05 Quanze Chen , Daniel S. Weld , Amy X. Zhang

Algorithmic information theory studies description complexity and randomness and is now a well known field of theoretical computer science and mathematical logic. There are several textbooks and monographs devoted to this theory where one…

Information Theory · Computer Science 2015-04-21 Alexander Shen

A method for selecting a graphical model for $p$-vector-valued stationary Gaussian time series was recently proposed by Matsuda and uses the Kullback-Leibler divergence measure to define a test statistic. This statistic was used in a…

Applications · Statistics 2023-07-19 R. J. Wolstenholme , A. T. Walden

We study distributed algorithms for some fundamental problems in data summarization. Given a communication graph $G$ of $n$ nodes each of which may hold a value initially, we focus on computing $\sum_{i=1}^N g(f_i)$, where $f_i$ is the…

Data Structures and Algorithms · Computer Science 2019-08-07 Hsin-Hao Su , Hoa T. Vu

We develop novel clustering algorithms for functional data when the number of clusters $K$ is unknown and also when it is prefixed. These algorithms are developed based on the Maximum Mean Discrepancy (MMD) measure between two sets of…

Methodology · Statistics 2025-07-16 Sourav Chakrabarty , Anirvan Chakraborty , Shyamal K. De

A number of engineering and scientific problems require representing and manipulating probability distributions over large alphabets, which we may think of as long vectors of reals summing to $1$. In some cases it is required to represent…

Information Theory · Computer Science 2023-10-27 Aviv Adler , Jennifer Tang , Yury Polyanskiy

In multi-center clinical trials, due to various reasons, the individual-level data are strictly restricted to be assessed publicly. Instead, the summarized information is widely available from published results. With the advance of…

Methodology · Statistics 2021-01-05 Jing Qin , Yukun Liu , Pengfei Li