English
Related papers

Related papers: Learning Classifiers with Fenchel-Young Losses: Ge…

200 papers

General probabilistic theories are designed to provide operationally the most general probabilistic models including both classical and quantum theories. In this letter, we introduce a systematic method to construct a series of entropies,…

Quantum Physics · Physics 2016-10-26 Gen Kimura , Junji Ishiguro , Makoto Fukui

We consider the generalization ability of algorithms for learning to rank at a query level, a problem also called subset ranking. Existing generalization error bounds necessarily degrade as the size of the document list associated with a…

Machine Learning · Computer Science 2016-08-24 Ambuj Tewari , Sougata Chaudhuri

We consider composite loss functions for multiclass prediction comprising a proper (i.e., Fisher-consistent) loss over probability distributions and an inverse link function. We establish conditions for their (strong) convexity and explore…

Machine Learning · Computer Science 2012-06-22 Mark Reid , Robert Williamson , Peng Sun

In this paper we provide a systematic exposition of basic properties of integrated distribution and quantile functions. We define these transforms in such a way that they characterize any probability distribution on the real line and are…

Probability · Mathematics 2018-01-04 Alexander A. Gushchin , Dmitriy A. Borzykh

We introduce a novel deep learning algorithm for computing convex conjugates of differentiable convex functions, a fundamental operation in convex analysis with various applications in different fields such as optimization, control theory,…

Machine Learning · Computer Science 2026-01-21 Aleksey Minabutdinov , Patrick Cheridito

In this paper, the method of gaps, a technique for deriving closed-form expressions in terms of information measures for the generalization error of supervised machine learning algorithms is introduced. The method relies on the notion of…

Machine Learning · Computer Science 2026-01-01 Samir M. Perlaza , Xinying Zou

Why do neural networks trained with large learning rates for a longer time often lead to better generalization? In this paper, we delve into this question by examining the relation between training and testing loss in neural networks.…

Machine Learning · Computer Science 2024-01-23 Yinuo Ren , Chao Ma , Lexing Ying

Listwise learning-to-rank methods form a powerful class of ranking algorithms that are widely adopted in applications such as information retrieval. These algorithms learn to rank a set of items by optimizing a loss that is a function of…

Machine Learning · Computer Science 2021-02-08 Sebastian Bruch

We give bounds on the difference between the weighted arithmetic mean and the weighted geometric mean. These imply refined Young inequalities and the reverses of the Young inequality. We also study some properties on the difference between…

Classical Analysis and ODEs · Mathematics 2021-04-28 Shigeru Furuichi , Nicuşor Minculete

Evolutionary computation can be used to optimize several different aspects of neural network architectures. For instance, the TaylorGLO method discovers novel, customized loss functions, resulting in improved performance, faster training,…

Machine Learning · Computer Science 2025-06-12 Santiago Gonzalez , Xin Qiu , Risto Miikkulainen

Empirical risk minimization (ERM) with a computationally feasible surrogate loss is a widely accepted approach for classification. Notably, the convexity and calibration (CC) properties of a loss function ensure consistency of ERM in…

Machine Learning · Statistics 2024-09-05 Ben Dai

The loss function is a key component in deep learning models. A commonly used loss function for classification is the cross entropy loss, which is a simple yet effective application of information theory for classification problems. Based…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Zeyu Song , Dongliang Chang , Zhanyu Ma , Xiaoxu Li , Zheng-Hua Tan

We investigate the in-distribution generalization of machine learning algorithms. We depart from traditional complexity-based approaches by analyzing information-theoretic bounds that quantify the dependence between a learning algorithm and…

Machine Learning · Statistics 2024-08-27 Borja Rodríguez-Gálvez , Ragnar Thobaben , Mikael Skoglund

Unsupervised representation learning methods are widely used for gaining insight into high-dimensional, unstructured, or structured data. In some cases, users may have prior topological knowledge about the data, such as a known cluster…

Machine Learning · Computer Science 2023-11-08 Edith Heiter , Robin Vandaele , Tijl De Bie , Yvan Saeys , Jefrey Lijffijt

We find the value of constants related to constraints in characterization of some known statistical distributions and then we proceed to use the idea behind maximum entropy principle to derive generalized version of this distributions using…

Statistical Mechanics · Physics 2007-05-23 Oscar Sotolongo-Costa , Alejandro Gonzalez Gonzalez , Francois Brouers

Continual learning (CL), which aims to learn a sequence of tasks, has attracted significant recent attention. However, most work has focused on the experimental performance of CL, and theoretical studies of CL are still limited. In…

Machine Learning · Computer Science 2023-02-14 Sen Lin , Peizhong Ju , Yingbin Liang , Ness Shroff

Generative neural networks learn how to produce highly realistic images from a large, but finite number of examples - or do they simply memorise their training set? To settle this question, Kadkhodaie, Guth, Simoncelli and Mallat (ICLR '24)…

Machine Learning · Statistics 2026-05-21 Antoine Maillard , Sebastian Goldt

In this study an attempt has been made to propose a way to develop new distribution. For this purpose, we need only idea about distribution function. Some important statistical properties of the new distribution like moments, cumulants,…

Methodology · Statistics 2024-08-30 Brijesh P. Singh , Utpal Dhar Das

Shannon's entropy and other entropy-based concepts are derived from the new, more general concept of relative divergence of one "grading' function on a linearly ordered set from another such function. The definition of relative divergence…

Probability · Mathematics 2019-03-14 Alexander Dukhovny

In information theory, one major goal is to find useful functions that summarize the amount of information contained in the interaction of several random variables. Specifically, one can ask how the classical Shannon entropy, mutual…

Information Theory · Computer Science 2025-02-14 Leon Lang , Pierre Baudot , Rick Quax , Patrick Forré