English
Related papers

Related papers: On Data-Processing and Majorization Inequalities f…

200 papers

Transformers exhibit a notable property of \emph{size generalization}, demonstrating an ability to extrapolate from smaller token sets to significantly longer ones. This behavior has been documented across diverse applications, including…

Machine Learning · Computer Science 2026-01-12 Anastasiia Alokhina , Pan Li

We consider a system composed of a fixed number of particles with total energy smaller or equal to some prescribed value. The particles are non-interacting, indistinguishable and distributed over fixed number of energy levels. The energy…

Probability · Mathematics 2021-03-23 Tomasz M. Łapiński

The quantum relative entropy is a measure of the distinguishability of two quantum states, and it is a unifying concept in quantum information theory: many information measures such as entropy, conditional entropy, mutual information, and…

Quantum Physics · Physics 2018-08-13 Mark M. Wilde

Data analysis and machine learning have become an integrative part of the modern scientific methodology, offering automated procedures for the prediction of a phenomenon based on past observations, unraveling underlying patterns in data and…

Machine Learning · Statistics 2015-06-04 Gilles Louppe

A data processing inequality states that the quantity of shared information between two entities (e.g. signals, strings) cannot be significantly increased when one of the entities is processed by certain kinds of transformations. In this…

Computational Complexity · Computer Science 2016-08-18 Adam Case

This paper presents a novel information-theoretic perspective on generalization in machine learning by framing the learning problem within the context of lossy compression and applying finite blocklength analysis. In our approach, the…

Machine Learning · Computer Science 2026-02-05 Kosuke Sugiyama , Masato Uchida

Can autoregressive large language models (LLMs) learn consistent probability distributions when trained on sequences in different token orders? We prove formally that for any well-defined probability distribution, sequence perplexity is…

Computation and Language · Computer Science 2025-05-14 Xiaoliang Luo , Xinyi Xu , Michael Ramscar , Bradley C. Love

Some new results are derived concerning random coding error exponents and expurgated exponents for list decoding with a deterministic list size $L$. Two asymptotic regimes are considered, the fixed list-size regime, where $L$ is fixed…

Information Theory · Computer Science 2016-11-17 Neri Merhav

Submodular function optimization has numerous applications in machine learning and data analysis, including data summarization which aims to identify a concise and diverse set of data points from a large dataset. It is important to…

Data Structures and Algorithms · Computer Science 2023-04-11 Shaojie Tang , Jing Yuan , Twumasi Mensah-Boateng

Drawing on an analogy with the second law of thermodynamics for adiabatically isolated systems, Cover argued that data-processing inequalities may be seen as second laws for "computationally isolated systems," namely, systems evolving…

Quantum Physics · Physics 2019-01-07 Francesco Buscemi

We develop a unified theory to analyze the microcanonical ensembles with several constraints given by unbounded observables. Several interesting phenomena that do not occur in the single constraint case can happen under the multiple…

Probability · Mathematics 2019-01-24 Kyeongsik Nam

Entropic uncertainty relations are interesting in their own rights as well as for a lot of applications. Keeping this in mind, we try to make the corresponding inequalities as tight as possible. The use of parametrized entropies also allows…

Quantum Physics · Physics 2023-05-30 Alexey E. Rastegin

Shannon based his information theory on the notion of probability measures as it we developed by Kolmogorov. In this paper we study some fundamental problems in information theory based on expectation measures. In the theory of expectation…

Information Theory · Computer Science 2025-01-30 Peter Harremoës

A variety of techniques have been proposed to train machine learning classifiers that are independent of a given feature. While this can be an essential technique for enabling background estimation, it may also be useful for reducing…

High Energy Physics - Phenomenology · Physics 2022-02-09 Aishik Ghosh , Benjamin Nachman

Inspired by coarse-graining approaches used in physics, we show how similar algorithms can be adapted for data. The resulting algorithms are based on layered tree tensor networks and scale linearly with both the dimension of the input and…

Machine Learning · Statistics 2018-05-01 E. M. Stoudenmire

Extremes of information combining inequalities play an important role in the analysis of sparse-graph codes under message-passing decoding. We introduce new tools for the derivation of such inequalities, and show by means of a concrete…

Information Theory · Computer Science 2012-02-01 Lucas Boczkowski

We introduce two new classes of measures of information for statistical experiments which generalise and subsume $\phi$-divergences, integral probability metrics, $\mathfrak{N}$-distances (MMD), and $(f,\Gamma)$ divergences between two or…

Machine Learning · Computer Science 2023-09-11 Robert C. Williamson , Zac Cranko

Despite widespread success in language understanding and generation, large language models (LLMs) exhibit unclear and often inconsistent behavior when faced with tasks that require probabilistic reasoning. In this work, we present the first…

Computation and Language · Computer Science 2025-09-29 Mobina Pournemat , Keivan Rezaei , Gaurang Sriramanan , Arman Zarei , Jiaxiang Fu , Yang Wang , Hamid Eghbalzadeh , Soheil Feizi

Recent literature in the last Maximum Entropy workshop introduced an analogy between cumulative probability distributions and normalized utility functions. Based on this analogy, a utility density function can de defined as the derivative…

Artificial Intelligence · Computer Science 2009-11-10 Ali E. Abbas

The method of optimizing entropy is used to (i) conduct Asymptotic Hypothesis Testing and (ii) determine the particle distribution for which Entropy is maximized. This paper focuses on two related applications of Information Theory:…

Statistics Theory · Mathematics 2016-03-09 Khizar Qureshi
‹ Prev 1 4 5 6 7 8 10 Next ›