English
Related papers

Related papers: Information Measures: the Curious Case of the Bina…

200 papers

Can autoregressive large language models (LLMs) learn consistent probability distributions when trained on sequences in different token orders? We prove formally that for any well-defined probability distribution, sequence perplexity is…

Computation and Language · Computer Science 2025-05-14 Xiaoliang Luo , Xinyi Xu , Michael Ramscar , Bradley C. Love

Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) reasoning degradation, (4) retrieval fragility, and (5)…

An alternative proof is given of the existence of greatest lower bounds in the imbalance order of binary maximal instantaneous codes of a given size. These codes are viewed as maximal antichains of a given size in the infinite binary tree…

Combinatorics · Mathematics 2017-10-09 Stephan Foldes , D. Stott Parker , Sandor Radeleczki

The index coding problem is studied from an interference alignment perspective, providing new results as well as new insights into, and generalizations of, previously known results. An equivalence is established between multiple unicast…

Information Theory · Computer Science 2012-05-08 Hamed Maleki , Viveck R. Cadambe , Syed A. Jafar

How should text dataset sizes be compared across languages? Even for content-matched (parallel) corpora, UTF-8 encoded text can require a dramatically different number of bytes for different languages. In our work, we define the byte…

Computation and Language · Computer Science 2024-03-04 Catherine Arnett , Tyler A. Chang , Benjamin K. Bergen

The properties of precessing, coalescing binary black holes are presently inferred through comparison with two approximate models of compact binary coalescence. In this work we show these two models often disagree substantially when…

General Relativity and Quantum Cosmology · Physics 2018-01-03 A. R. Williamson , J. Lange , R. O'Shaughnessy , J. A. Clark , P. Kumar , J. Calderón Bustillo , J. Veitch

The $k$-Means clustering problem on $n$ points is NP-Hard for any dimension $d\ge 2$, however, for the 1D case there exists exact polynomial time algorithms. Previous literature reported an $O(kn^2)$ time dynamic programming algorithm that…

Data Structures and Algorithms · Computer Science 2018-04-26 Allan Grønlund , Kasper Green Larsen , Alexander Mathiasen , Jesper Sindahl Nielsen , Stefan Schneider , Mingzhou Song

We introduce a family of information leakage measures called maximal $\alpha,\beta$-leakage, parameterized by real numbers $\alpha$ and $\beta$. The measure is formalized via an operational definition involving an adversary guessing an…

Information Theory · Computer Science 2022-11-29 Atefeh Gilani , Gowtham R. Kurri , Oliver Kosut , Lalitha Sankar

A deterministic finite automaton (DFA) separates two strings $w$ and $x$ if it accepts $w$ and rejects $x$. The minimum number of states required for a DFA to separate $w$ and $x$ is denoted by $sep(w,x)$. The present paper shows that the…

Formal Languages and Automata Theory · Computer Science 2018-02-13 Farzam Ebrahimnejad

Effective complexity measures the information content of the regularities of an object. It has been introduced by M. Gell-Mann and S. Lloyd to avoid some of the disadvantages of Kolmogorov complexity, also known as algorithmic information…

Information Theory · Computer Science 2010-11-22 Nihat Ay , Markus Mueller , Arleta Szkola

The Bregman divergence have been the subject of several studies. We do not go to do an exhaustive study of its subclasses, but propose a proof that shows that the \b{eta}-divergence are subclasses of the Bregman divergences. It is in this…

Methodology · Statistics 2018-05-21 Macoumba Ndourand Mactar Ndaw , Papa Ngom

The classification of random objects within metric spaces without a vector structure has attracted increasing attention. However, the complexity inherent in such non-Euclidean data often restricts existing models to handle only a limited…

Methodology · Statistics 2024-03-20 Shuaida He , Jiaqi Li , Xin Chen

Mixed $f$-divergences, a concept from information theory and statistics, measure the difference between multiple pairs of distributions. We introduce them for log concave functions and establish some of their properties. Among them are…

Functional Analysis · Mathematics 2016-06-29 Umut Caglar , Elisabeth M. Werner

The binary divergences that are divergences between probability measures defined on the same 2-point set have an interesting property. For the chi-squared divergence and the relative entropy, it is known that their binary divergence attain…

Information Theory · Computer Science 2021-02-09 Tomohiro Nishiyama

Advancements in Large Language Models (LLMs) have increased the performance of different natural language understanding as well as generation tasks. Although LLMs have breached the state-of-the-art performance in various tasks, they often…

Computation and Language · Computer Science 2025-05-28 Charaka Vinayak Kumar , Ashok Urlana , Gopichand Kanumolu , Bala Mallikarjunarao Garlapati , Pruthwik Mishra

Discrepancy measures how uniformly distributed a point set is with respect to a given set of ranges. There are two notions of discrepancy, namely continuous discrepancy and combinatorial discrepancy. Depending on the ranges, several…

Computational Geometry · Computer Science 2011-03-24 Panos Giannopoulos , Christian Knauer , Magnus Wahlström , Daniel Werner

This paper studies the mismatched decoding problem for binary-input discrete memoryless channels. An example is provided for which an achievable rate based on superposition coding exceeds the LM rate (Hui, 1983; Csisz\'ar-K\"orner, 1981),…

Information Theory · Computer Science 2015-08-11 Jonathan Scarlett , Anelia Somekh-Baruch , Alfonso Martinez , Albert Guillén i Fàbregas

Large Language Models (LLMs) are widely deployed in real-world applications, yet little is known about their training dynamics at the token level. Evaluation typically relies on aggregated training loss, measured at the batch level, which…

Computation and Language · Computer Science 2024-10-17 Andrea Pinto , Tomer Galanti , Randall Balestriero

Unmeasured confounding is a threat to causal inference and gives rise to biased estimates. In this article, we consider the problem of individualized decision-making under partial identification. Firstly, we argue that when faced with…

Methodology · Statistics 2021-10-22 Yifan Cui

This work provides data-processing and majorization inequalities for $f$-divergences, and it considers some of their applications to coding problems. This work also provides tight bounds on the R\'{e}nyi entropy of a function of a discrete…

Information Theory · Computer Science 2021-04-01 Igal Sason