English
Related papers

Related papers: Compression for Quadratic Similarity Queries

200 papers

We study linear chance-constrained problems where the coefficients follow a Gaussian mixture distribution. We provide mixed-binary quadratic programs that give inner and outer approximations of the chance constraint based on piecewise…

Optimization and Control · Mathematics 2025-11-24 Shibshankar Dey , Sanjay Mehrotra , Anirudh Subramanyam

In practical data integration systems, it is common for the data sources being integrated to provide conflicting information about the same entity. Consequently, a major challenge for data integration is to derive the most complete and…

Databases · Computer Science 2012-03-05 Bo Zhao , Benjamin I. P. Rubinstein , Jim Gemmell , Jiawei Han

We introduce a new distance-preserving compact representation of multi-dimensional point-sets. Given $n$ points in a $d$-dimensional space where each coordinate is represented using $B$ bits (i.e., $dB$ bits per point), it produces a…

Data Structures and Algorithms · Computer Science 2017-11-07 Piotr Indyk , Ilya Razenshteyn , Tal Wagner

We propose a new \textit{quadratic programming-based} method of approximating a nonstandard density using a multivariate Gaussian density. Such nonstandard densities usually arise while developing posterior samplers for unobserved…

Econometrics · Economics 2023-02-14 Abhishek K. Umrawal , Joshua C. C. Chan

In various data settings, it is necessary to compare observations from disparate data sources. We assume the data is in the dissimilarity representation and investigate a joint embedding method that results in a commensurate representation…

Methodology · Statistics 2016-01-05 Sancar Adali , Carey E. Priebe

Random Number Generators play a critical role in a number of important applications. In practice, statistical testing is employed to gather evidence that a generator indeed produces numbers that appear to be random. In this paper, we…

Computational Complexity · Computer Science 2010-03-25 Weiling Chang , Binxing Fang , Xiaochun Yun , Shupeng Wang , Xiangzhan Yu

Today's exponentially increasing data volumes and the high cost of storage make compression essential for the Big Data industry. Although research has concentrated on efficient compression, fast decompression is critical for analytics…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-06-03 Evangelia Sitaridi , Rene Mueller , Tim Kaldewey , Guy Lohman , Kenneth Ross

Data deduplication saves storage space by identifying and removing repeats in the data stream. Compared with traditional compression methods, data deduplication schemes are more time efficient and are thus widely used in large scale storage…

Information Theory · Computer Science 2022-05-30 Hao Lou , Farzad Farnoud

The information spectrum approach gives general formulae for optimal rates of various information theoretic protocols, under minimal assumptions on the nature of the sources, channels and entanglement resources involved. This paper…

Quantum Physics · Physics 2007-05-23 Garry Bowen , Nilanjana Datta

In Sequential Massive Random Access (SMRA), a set of correlated sources is jointly encoded and stored on a server, and clients want to access to only a subset of the sources. Since the number of simultaneous clients can be huge, the server…

Information Theory · Computer Science 2018-01-18 Elsa Dupraz , Thomas Maugey , Aline Roumy , Michel Kieffer

We provide a framework for one-shot quantum rate distortion coding, in which the goal is to determine the minimum number of qubits required to compress quantum information as a function of the probability that the distortion incurred upon…

Quantum Physics · Physics 2013-12-09 Nilanjana Datta , Joseph M. Renes , Renato Renner , Mark M. Wilde

The distortion-rate function of output-constrained lossy source coding with limited common randomness is analyzed for the special case of squared error distortion measure. An explicit expression is obtained when both source and…

Information Theory · Computer Science 2024-03-25 Li Xie , Liangyan Li , Jun Chen , Zhongshan Zhang

Data compression algorithms typically rely on identifying repeated sequences of symbols from the original data to provide a compact representation of the same information, while maintaining the ability to recover the original data from the…

Databases · Computer Science 2023-08-08 Francesco Taurone , Daniel E. Lucani , Marcell Fehér , Qi Zhang

In this work, we employ the Bayesian inference framework to solve the problem of estimating the solution and particularly, its derivatives, which satisfy a known differential equation, from the given noisy and scarce observations of the…

Computation · Statistics 2020-10-09 Hongqiao Wang , Xiang Zhou

Understanding the impact of data structure on the computational tractability of learning is a key challenge for the theory of neural networks. Many theoretical works do not explicitly model training data, or assume that inputs are drawn…

Machine Learning · Statistics 2022-05-23 Sebastian Goldt , Bruno Loureiro , Galen Reeves , Florent Krzakala , Marc Mézard , Lenka Zdeborová

We revisit the replica method for analyzing inference and learning in parametric models, considering situations where the data-generating distribution is unknown or analytically intractable. Instead of assuming idealized distributions to…

Disordered Systems and Neural Networks · Physics 2025-11-17 Takashi Takahashi

Computable Information Density (CID), the ratio of the length of a losslessly compressed data file to that of the uncompressed file, is a measure of order and correlation in both equilibrium and nonequilibrium systems. Here we show that…

Statistical Mechanics · Physics 2020-10-26 Stefano Martiniani , Yuval Lemberg , Paul M. Chaikin , Dov Levine

Approximating the solution of the nonlinear filtering problem with Gaussian mixtures has been a very popular method since the 1970s. However, the vast majority of such approximations are introduced in an ad-hoc manner without theoretical…

Probability · Mathematics 2014-01-28 Dan Crisan , Kai Li

In recent years, as data and problem sizes have increased, distributed learning has become an essential tool for training high-performance models. However, the communication bottleneck, especially for high-dimensional data, is a challenge.…

Optimization and Control · Mathematics 2025-04-28 Dmitry Bylinkin , Aleksandr Beznosikov

This paper studies the minimum achievable source coding rate as a function of blocklength $n$ and probability $\epsilon$ that the distortion exceeds a given level $d$. Tight general achievability and converse bounds are derived that hold at…

Information Theory · Computer Science 2016-11-15 Victoria Kostina , Sergio Verdú
‹ Prev 1 8 9 10 Next ›