English
Related papers

Related papers: On information gain, Kullback-Leibler divergence, …

200 papers

Kullback-Leibler (KL) divergence is a fundamental concept in information theory that quantifies the discrepancy between two probability distributions. In the context of Variational Autoencoders (VAEs), it serves as a central regularization…

Machine Learning · Computer Science 2026-04-14 Andrés Muñoz , Rodrigo Ramele

This work proposed kernel selection approaches for probabilistic classifiers based on features produced by the convolutional encoder of a variational autoencoder. Particularly, the developed methodologies allow the selection of the most…

Machine Learning · Computer Science 2025-08-05 Fábio Mendonça , Sheikh Shanawaz Mostafa , Fernando Morgado-Dias , Antonio G. Ravelo-García

Transfer learning, or domain adaptation, is concerned with machine learning problems in which training and testing data come from possibly different probability distributions. In this work, we give an information-theoretic analysis of the…

Information Theory · Computer Science 2024-08-09 Xuetong Wu , Jonathan H. Manton , Uwe Aickelin , Jingge Zhu

We show that the moment generating function of the Kullback-Leibler divergence (relative entropy) between the empirical distribution of $n$ independent samples from a distribution $P$ over a finite alphabet of size $k$ (i.e. a multinomial…

Information Theory · Computer Science 2020-10-06 Rohit Agrawal

Predictive inference requires balancing statistical accuracy against informational complexity, yet the choice of complexity measure is usually imposed rather than derived. We treat econometric objects as predictive rules, mappings from…

Statistics Theory · Mathematics 2026-02-16 Nicholas G. Polson , Daniel Zantedeschi

Mutual information $I(X;Y)$ is a useful definition in information theory to estimate how much information the random variable $Y$ holds about the random variable $X$. One way to define the mutual information is by comparing the joint…

Information Theory · Computer Science 2022-04-14 Bulut Kuskonmaz , Jaron Skovsted Gundersen , Rafal Wisniewski

New concepts from nonequilibrium thermodynamics are used to show that Landauer's principle can be understood in terms of time asymmetry in the dynamical randomness generated by the physical process of the erasure of digital information. In…

Statistical Mechanics · Physics 2009-11-13 D. Andrieux , P. Gaspard

In this work we present a new method for the estimation of Mutual Information (MI) between random variables. Our approach is based on an original interpretation of the Girsanov theorem, which allows us to use score-based diffusion models to…

Machine Learning · Computer Science 2024-05-16 Giulio Franzese , Mustapha Bounoua , Pietro Michiardi

Empirical data can often be considered as samples from a set of probability distributions. Kernel methods have emerged as a natural approach for learning to classify these distributions. Although numerous kernels between distributions have…

Machine Learning · Computer Science 2024-12-02 Oleksii Kachaiev , Stefano Recanatesi

We characterize Martin-L\"of randomness and Schnorr randomness in terms of the merging of opinions, along the lines of the Blackwell-Dubins Theorem. After setting up a general framework for defining notions of merging randomness, we focus…

Logic · Mathematics 2026-03-10 Simon M. Huttegger , Sean Walsh , Francesca Zaffora Blando

Accurately matching visual and textual data in cross-modal retrieval has been widely studied in the multimedia community. To address these challenges posited by the heterogeneity gap and the semantic gap, we propose integrating Shannon…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Wei Chen , Yu Liu , Erwin M. Bakker , Michael S. Lew

The information content of a source is defined in terms of the minimum number of bits needed to store the output of the source in a perfectly recoverable way. A similar definition can be given in the case of quantum sources, with qubits…

Quantum Physics · Physics 2023-04-03 Paolo Perinotti , Alessandro Tosini , Leonardo Vaglini

The measure-theoretic definition of Kullback-Leibler relative-entropy (KL-entropy) plays a basic role in the definitions of classical information measures. Entropy, mutual information and conditional forms of entropy can be expressed in…

Mathematical Physics · Physics 2007-05-23 Ambedkar Dukkipati , Shalabh Bhatnagar , M Narasimha Murty

We introduce the loss kernel, an interpretability method for measuring similarity between data points according to a trained neural network. The kernel is the covariance matrix of per-sample losses computed under a distribution of…

Machine Learning · Computer Science 2025-10-01 Maxwell Adam , Zach Furman , Jesse Hoogland

With the exponential growth in data volume and the emergence of data-intensive applications, particularly in the field of machine learning, concerns related to resource utilization, privacy, and fairness have become paramount. This paper…

Computation and Language · Computer Science 2024-05-14 Kaan Kale , Homa Esfahanizadeh , Noel Elias , Oguzhan Baser , Muriel Medard , Sriram Vishwanath

Modern machine learning approaches excel in static settings where a large amount of i.i.d. training data are available for a given task. In a dynamic environment, though, an intelligent agent needs to be able to transfer knowledge and…

Machine Learning · Computer Science 2023-03-13 Jonas Wildberger , Siyuan Guo , Arnab Bhattacharyya , Bernhard Schölkopf

R\'enyi divergence is related to R\'enyi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by R\'enyi as a measure of information that satisfies almost the same…

Information Theory · Computer Science 2014-04-25 Tim van Erven , Peter Harremoës

After Shannon, entropy becomes a fundamental quantity to describe not only uncertainity or chaos of a system but also information carried by the system. Shannon's important discovery is to give a mathematical expression of the mutual…

Quantum Physics · Physics 2007-05-23 Masanori Ohya

Knowledge distillation involves transferring knowledge from large, cumbersome teacher models to more compact student models. The standard approach minimizes the Kullback-Leibler (KL) divergence between the probabilistic outputs of a teacher…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Nikolaos Giakoumoglou , Tania Stathaki

This paper investigates the accuracy of generative models and the impact of knowledge transfer on their generation precision. Specifically, we examine a generative model for a target task, fine-tuned using a pre-trained model from a source…

Machine Learning · Statistics 2025-06-03 Xinyu Tian , Xiaotong Shen