English
Related papers

Related papers: Dimension-free Information Concentration via Exp-C…

200 papers

Maximum entropy models provide the least constrained probability distributions that reproduce statistical properties of experimental datasets. In this work we characterize the learning dynamics that maximizes the log-likelihood in the case…

Disordered Systems and Neural Networks · Physics 2016-09-21 Ulisse Ferrari

The information bottleneck framework provides a systematic approach to learning representations that compress nuisance information in the input and extract semantically meaningful information about predictions. However, the choice of a…

Gathering the most information by picking the least amount of data is a common task in experimental design or when exploring an unknown environment in reinforcement learning and robotics. A widely used measure for quantifying the…

Machine Learning · Statistics 2015-09-17 Johannes Kulick , Robert Lieck , Marc Toussaint

Exponential models of distributions are widely used in machine learning for classiffication and modelling. It is well known that they can be interpreted as maximum entropy models under empirical expectation constraints. In this work, we…

Machine Learning · Computer Science 2012-07-19 Amir Globerson , Naftali Tishby

In this paper, we study two problems: (1) estimation of a $d$-dimensional log-concave distribution and (2) bounded multivariate convex regression with random design with an underlying log-concave density or a compactly supported…

Statistics Theory · Mathematics 2020-02-21 Gil Kur , Yuval Dagan , Alexander Rakhlin

We study the source of uncertainty in DeepSeek R1-32B by analyzing its self-reported verbal confidence on question answering (QA) tasks. In the default answer-then-confidence setting, the model is regularly over-confident, whereas semantic…

Computation and Language · Computer Science 2025-11-06 Jakub Podolak , Rajeev Verma

We consider a generic class of log-concave, possibly random, (Gibbs) measures. We prove the concentration of an infinite family of order parameters called multioverlaps. Because they completely parametrise the quenched Gibbs measure of the…

Probability · Mathematics 2022-12-22 Jean Barbier , Dmitry Panchenko , Manuel Sáenz

As conventional communication systems based on classic information theory have closely approached the limits of Shannon channel capacity, semantic communication has been recognized as a key enabling technology for the further improvement of…

Information Theory · Computer Science 2023-06-06 Jiancheng Tang , Qianqian Yang , Zhaoyang Zhang

We show that a language model's ability to predict text is tightly linked to the breadth of its embedding space: models that spread their contextual representations more widely tend to achieve lower perplexity. Concretely, we find that…

Computation and Language · Computer Science 2026-04-21 Yanhong Li , Ming Li , Karen Livescu , Jiawei Zhou

Directed information or its variants are utilized extensively in the characterization of the capacity of channels with memory and feedback, nonanticipative lossy data compression, and their generalizations to networks. In this paper, we…

Information Theory · Computer Science 2015-12-24 Charalambos D. Charalambous , Photios A. Stavrou

We pedagogically present the information theory as originally established, explaining its essential ideas and paying attention to the expression employed to measure the amount of information. Also we discussed relationships between…

Quantum Physics · Physics 2019-12-10 Wallas S. Nascimento , Marcos M. de Almeida , Frederico V. Prudente

Scaling large language models (LLMs) leads to an emergent capacity to learn in-context from example demonstrations. Despite progress, theoretical understanding of this phenomenon remains limited. We argue that in-context learning relies on…

Computation and Language · Computer Science 2023-03-15 Michael Hahn , Navin Goyal

Recently, self-supervised learning has attracted great attention, since it only requires unlabeled data for model training. Contrastive learning is one popular method for self-supervised learning and has achieved promising empirical…

Machine Learning · Computer Science 2023-03-03 Weiran Huang , Mingyang Yi , Xuyang Zhao , Zihao Jiang

In high-dimensional problems, choosing a prior distribution such that the corresponding posterior has desirable practical and theoretical properties can be challenging. This begs the question: can the data be used to help choose a good…

Statistics Theory · Mathematics 2019-09-25 Ryan Martin , Stephen G. Walker

We develop a unified analysis of how information captures attention. A decision maker (DM) faces a dynamic information structure and decides when to stop paying attention. We characterize the convex$\unicode{x2013}$order frontier and…

Theoretical Economics · Economics 2024-09-25 Andrew Koh , Sivakorn Sanguanmoo

Information estimates such as the ``direct method'' of Strong et al. (1998) sidestep the difficult problem of estimating the joint distribution of response and stimulus by instead estimating the difference between the marginal and…

Neurons and Cognition · Quantitative Biology 2008-07-19 Vincent Q. Vu , Bin Yu , Robert E. Kass

We study the problem of maximum likelihood estimation of densities that are log-concave and lie in the graphical model corresponding to a given undirected graph $G$. We show that the maximum likelihood estimate (MLE) is the product of the…

Statistics Theory · Mathematics 2025-12-02 Kaie Kubjas , Olga Kuznetsova , Elina Robeva , Pardis Semnani , Luca Sodomaco

We generalize a result by Carlen and Cordero-Erausquin on the equivalence between the Brascamp-Lieb inequality and the subadditivity of relative entropy by allowing for random transformations (a broadcast channel). This leads to a unified…

Information Theory · Computer Science 2016-05-11 Jingbo Liu , Thomas A. Courtade , Paul Cuff , Sergio Verdu

Recently, self-supervised contrastive learning has achieved great success on various tasks. However, its underlying working mechanism is yet unclear. In this paper, we first provide the tightest bounds based on the widely adopted assumption…

Machine Learning · Computer Science 2025-11-06 Qi Zhang , Yifei Wang , Yisen Wang

This paper provides an elementary, self-contained analysis of diffusion-based sampling methods for generative modeling. In contrast to existing approaches that rely on continuous-time processes and then discretize, our treatment works…

Machine Learning · Statistics 2025-06-25 Galen Reeves , Henry D. Pfister
‹ Prev 1 4 5 6 7 8 10 Next ›