English
Related papers

Related papers: Correlation lengths in the language of computable …

200 papers

The Shannon entropy of a random variable $X$ has much behaviour analogous to a signed measure. Previous work has concretized this connection by defining a signed measure $\mu$ on an abstract information space $\tilde{X}$, which is taken to…

Information Theory · Computer Science 2023-05-15 Keenan J. A. Down , Pedro A. M. Mediano

Independent Component Analysis (ICA) offers interpretable semantic components of embeddings. While ICA theory assumes that embeddings can be linearly decomposed into independent components, real-world data often do not satisfy this…

Computation and Language · Computer Science 2024-10-10 Momose Oyama , Hiroaki Yamagiwa , Hidetoshi Shimodaira

Suppose that we are given a string $s$ of length $n$ over an alphabet $\{0,1,\ldots,n^{O(1)}\}$ and $\delta$ is the string complexity of $s$, a known compression measure. We describe an index on $s$ with $O(\delta\log\frac{n}{\delta})$…

Data Structures and Algorithms · Computer Science 2026-04-15 Dmitry Kosolobov

Information scrambling refers to the rapid spreading of initially localized information over an entire system, via the generation of global entanglement. This effect is usually detected by measuring a temporal decay of the out-of-time order…

Quantum Physics · Physics 2022-07-28 Joseph Harris , Bin Yan , Nikolai A. Sinitsyn

The shapes of galaxy N-point correlation functions can be used as standard rulers to constrain the distance-redshift relationship and thence the expansion rate of the Universe. The cosmological density fields traced by late-time galaxy…

Cosmology and Nongalactic Astrophysics · Physics 2021-07-07 Lado Samushia , Zachary Slepian , Francisco Villaescusa-Navarro

We present a new information-theoretic definition and associated results, based on list decoding in a source coding setting. We begin by presenting list-source codes, which naturally map a key length (entropy) to list size. We then show…

Information Theory · Computer Science 2012-10-09 Flavio du Pin Calmon , Muriel Médard , Linda M. Zeger , João Barros , Mark M. Christiansen , Ken. R. Duffy

Understanding associations between paired high-dimensional longitudinal datasets is a fundamental yet challenging problem that arises across scientific domains, including longitudinal multi-omic studies. The difficulty stems from the…

Methodology · Statistics 2026-01-21 Jianbin Tan , Pixu Shi

The concept of information has emerged as a language in its own right, bridging several disciplines that analyze natural phenomena and man-made systems. Integrated information has been introduced as a metric to quantify the amount of…

Neurons and Cognition · Quantitative Biology 2019-06-10 Alberto Hernández-Espinosa , Héctor Zenil , Narsis A. Kiani , Jesper Tegnér

We describe the DISC (Different Individuals, Same Clusters) design, a sampling scheme that can improve the precision of difference-in-differences (DID) estimators in settings involving repeated sampling of a population at multiple time…

Methodology · Statistics 2025-08-21 Jordan Downey , Avi Kenny

Image compression has been a frequent topic of presentations at ADASS. Compression is often viewed as just a technique to fit more data into a smaller space. Rather, the packing of data - its "density" - affects every facet of local data…

Instrumentation and Methods for Astrophysics · Physics 2009-10-21 Robert L. Seaman , Richard L. White , William D. Pence

We present some new results which relate information to chaotic dynamics. In our approach the quantity of information is measured by the Algorithmic Information Content (Kolmogorov complexity) or by a sort of computable version of it…

Statistical Mechanics · Physics 2007-05-23 V. Benci , C. Bonanno , S. Galatolo , G. Menconi , M. Virgilio

Correlation clustering is a flexible framework for partitioning data based solely on pairwise similarity or dissimilarity information, without requiring the number of clusters as input. However, in many practical scenarios, these pairwise…

Machine Learning · Computer Science 2025-12-11 Linus Aronsson , Morteza Haghir Chehreghani

Dimensionality reduction is a crucial technique in data analysis, as it allows for the efficient visualization and understanding of high-dimensional datasets. The circular coordinate is one of the topological data analysis techniques…

Algebraic Topology · Mathematics 2023-01-31 Taejin Paik , Jaemin Park

A new family of codes, called clustering-correcting codes, is presented in this paper. This family of codes is motivated by the special structure of data that is stored in DNA-based storage systems. The data stored in these systems has the…

Information Theory · Computer Science 2019-03-12 Tal Shinkar , Eitan Yaakobi , Andreas Lenz , Antonia Wachter-Zeh

Configurational entropy (CE) and configurational complexity (CC) are recently popularized information theoretic measures used to study the stability of solitons. This paper examines their behavior for 2D and 3D lattice Ising Models, where…

Statistical Mechanics · Physics 2025-03-06 Damian R Sowinski , Sean Kelty , Gourab Ghoshal

Finding desired information from large data set is a difficult problem. Information retrieval is concerned with the structure, analysis, organization, storage, searching, and retrieval of information. Index is the main constituent of an IR…

Information Retrieval · Computer Science 2012-09-26 Md. Abdullah al Mamun , Md. Hanif , Md. Rakib Uddin , Tanvir Ahmed , Md. Mofizul Islam

Partial Information Decomposition (PID) has become one of the most prominent information-theoretic frameworks for describing the structure and quality of information in complex systems. Despite its widespread utility, there exists no unique…

Information Theory · Computer Science 2026-03-10 Alberto Liardi , Keenan J. A. Down , George Blackburne , Matteo Neri , Pedro A. M. Mediano

We study a generalization of deduplication, which enables lossless deduplication of highly similar data and show that standard deduplication with fixed chunk length is a special case. We provide bounds on the expected length of coded…

Information Theory · Computer Science 2020-03-04 Rasmus Vestergaard , Qi Zhang , Daniel E. Lucani

The improved method of intermittent data analysis is proposed. It exploits, in addition to the standard density moments, the information on the bin-bin correlations, observed in the data and expressed in terms of the density correlators.…

High Energy Physics - Phenomenology · Physics 2007-05-23 B. Ziaja

Lossy image coding standards such as JPEG and MPEG have successfully achieved high compression rates for human consumption of multimedia data. However, with the increasing prevalence of IoT devices, drones, and self-driving cars, machines…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Chen-Hsiu Huang , Ja-Ling Wu