English
Related papers

Related papers: Learning Classifiers with Fenchel-Young Losses: Ge…

200 papers

A generalized constitutive relation error is proposed in an analogous form to Fenchel-Young inequality on the basis of the key idea of Legendre-Fenchel duality theory. The generalized constitutive relation error is linked with the global…

Numerical Analysis · Mathematics 2016-11-18 Mengwu Guo , Weimin Han , Hongzhi Zhong

Convolutional Neural Networks (CNNs) are widely used for image classification in a variety of fields, including medical imaging. While most studies deploy cross-entropy as the loss function in such tasks, a growing number of approaches have…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Vasileios Baltatzis , Loic Le Folgoc , Sam Ellis , Octavio E. Martinez Manzanera , Kyriaki-Margarita Bintsi , Arjun Nair , Sujal Desai , Ben Glocker , Julia A. Schnabel

Since the celebrated works of Russo and Zou (2016,2019) and Xu and Raginsky (2017), it has been well known that the generalization error of supervised learning algorithms can be bounded in terms of the mutual information between their input…

Machine Learning · Statistics 2022-07-20 Gábor Lugosi , Gergely Neu

While machine learning (ML) architectures have evolved rapidly to account for complex data, loss functions like cross-entropy remain mostly structure-agnostic in many real-world applications. However, the `class-symmetric' nature of these…

Machine Learning · Computer Science 2026-05-28 Yasser Taha , Grégoire Montavon , Nils Körber

Learning with a {\it convex loss} function has been a dominating paradigm for many years. It remains an interesting question how non-convex loss functions help improve the generalization of learning with broad applicability. In this paper,…

Machine Learning · Computer Science 2018-05-22 Yi Xu , Shenghuo Zhu , Sen Yang , Chi Zhang , Rong Jin , Tianbao Yang

As shown in recent research, deep neural networks can perfectly fit randomly labeled data, but with very poor accuracy on held out data. This phenomenon indicates that loss functions such as cross-entropy are not a reliable indicator of…

Machine Learning · Statistics 2019-06-13 Yiding Jiang , Dilip Krishnan , Hossein Mobahi , Samy Bengio

Significant advances have been made recently on training neural networks, where the main challenge is in solving an optimization problem with abundant critical points. However, existing approaches to address this issue crucially rely on a…

Machine Learning · Computer Science 2019-02-28 Weihao Gao , Ashok Vardhan Makkuva , Sewoong Oh , Pramod Viswanath

The increased uncertainty and complexity of nonlinear systems have motivated investigators to consider generalized approaches to defining an entropy function. New insights are achieved by defining the average uncertainty in the probability…

Statistics Theory · Mathematics 2025-11-25 Kenric P. Nelson , Sabir Umarov , Mark A. Kon

The Softmax loss is one of the most widely employed surrogate objectives for classification and ranking tasks. To elucidate its theoretical properties, the Fenchel-Young framework situates it as a canonical instance within a broad family of…

Machine Learning · Computer Science 2026-02-02 Yuanhao Pu , Defu Lian , Enhong Chen

In this paper, we propose a new discriminative model named \emph{nonextensive information theoretical machine (NITM)} based on nonextensive generalization of Shannon information theory. In NITM, weight parameters are treated as random…

Machine Learning · Computer Science 2016-04-22 Chaobing Song , Shu-Tao Xia

We study prediction and estimation problems using empirical risk minimization, relative to a general convex loss function. We obtain sharp error rates even when concentration is false or is very restricted, for example, in heavy-tailed…

Machine Learning · Statistics 2014-10-14 Shahar Mendelson

We introduce a variational algorithm based on Matrix Product States that is trained by minimizing a generalized free energy defined using Tsallis entropy instead of the standard Gibbs entropy. As a result, our model can generate the…

Statistical Mechanics · Physics 2024-09-16 Pablo Díez-Valle , Fernando Martínez-García , Juan José García-Ripoll , Diego Porras

This paper studies the generalization performance of multi-class classification algorithms, for which we obtain, for the first time, a data-dependent generalization error bound with a logarithmic dependence on the class size, substantially…

Machine Learning · Computer Science 2015-06-16 Yunwen Lei , Ürün Dogan , Alexander Binder , Marius Kloft

The Perturbed Utility Model (PUM) framework provides a generalization of discrete choice analysis, unifying models like Multinomial Logit (MNL) and Sparsemax through convex optimization. However, standard Maximum Likelihood Estimation (MLE)…

Optimization and Control · Mathematics 2026-05-15 Xi Lin , Yafeng Yin , Tianming Liu

The logistic loss (a.k.a. cross-entropy loss) is one of the most popular loss functions used for multiclass classification. It is also the loss function of choice for next-token prediction in language modeling. It is associated with the…

Machine Learning · Computer Science 2025-06-16 Vincent Roulet , Tianlin Liu , Nino Vieillard , Michael E. Sander , Mathieu Blondel

The purpose of this note is to give the general solution of two functional equations connected to the Shannon entropy and also to the Tsallis entropy. As a result of this, we present the regular solution of these equations, as well.…

Classical Analysis and ODEs · Mathematics 2013-07-03 Eszter Gselmann

We give several inequalities on generalized entropies involving Tsallis entropies, using some inequalities obtained by improvements of Young's inequality. We also give a generalized Han's inequality.

Classical Analysis and ODEs · Mathematics 2013-01-08 S. Furuichi , N. Minculete , F. -C. Mitroi

Ensemble algorithms offer state of the art performance in many machine learning applications. A common explanation for their excellent performance is due to the bias-variance decomposition of the mean squared error which shows that the…

Machine Learning · Computer Science 2020-12-10 Sebastian Buschjäger , Lukas Pfahler , Katharina Morik

Researches using margin based comparison loss demonstrate the effectiveness of penalizing the distance between face feature and their corresponding class centers. Despite their popularity and excellent performance, they do not explicitly…

Computer Vision and Pattern Recognition · Computer Science 2020-06-12 Ying Huang , Shangfeng Qiu , Wenwei Zhang , Xianghui Luo , Jinzhuo Wang

Assuming that the loss function is convex in the prediction, we construct a prediction strategy universal for the class of Markov prediction strategies, not necessarily continuous. Allowing randomization, we remove the requirement of…

Machine Learning · Computer Science 2007-05-23 Vladimir Vovk