English
Related papers

Related papers: Generalized Mixability via Entropic Duality

200 papers

One of the central challenges in modern machine learning is understanding how neural networks generalize knowledge learned from training data to unseen test data. While numerous empirical techniques have been proposed to improve…

Machine Learning · Computer Science 2025-04-18 Entao Yang , Xiaotian Zhang , Yue Shang , Ge Zhang

In this paper, I expand Shannon's definition of entropy into a new form of entropy that allows integration of information from different random events. Shannon's notion of entropy is a special case of my more general definition of entropy.…

Machine Learning · Computer Science 2008-11-04 Stefan Jaeger

We study the connection between mixing properties for bipartite graphs and materialization of the mutual information in one-shot settings. We show that mixing properties of a graph imply impossibility to extract the mutual information…

Information Theory · Computer Science 2025-09-10 Geoffroy Caillat-Grenier , Andrei Romashchenko , Rustam Zyavgarov

A partial differential equation governing the global evolution of the joint probability distribution of an arbitrary number of local flow observations, drawn randomly from a control volume, is derived and applied to examples involving…

Fluid Dynamics · Physics 2026-01-14 John Craske , Paul Mannix

We consider a two-parameter family of random substitutions and show certain combinatorial and topological properties they satisfy. We establish that they admit recognisable words at every level. As a consequence, we get that the subshifts…

Dynamical Systems · Mathematics 2021-08-13 Giovanni B. Escolano , Neil Mañibo , Eden Delight Miro

Properness for supervised losses stipulates that the loss function shapes the learning algorithm towards the true posterior of the data generating distribution. Unfortunately, data in modern machine learning can be corrupted or twisted in…

Machine Learning · Computer Science 2022-02-02 Tyler Sypherd , Richard Nock , Lalitha Sankar

An information-theoretic upper bound on the generalization error of supervised learning algorithms is derived. The bound is constructed in terms of the mutual information between each individual training sample and the output of the…

Machine Learning · Computer Science 2020-08-06 Yuheng Bu , Shaofeng Zou , Venugopal V. Veeravalli

The Gaussian mixture model is widely used in unsupervised learning, owing to its simplicity and interpretability. However, a fundamental limitation of the classical Gaussian mixture model is that it forces each observation to belong to…

Machine Learning · Statistics 2026-04-27 Huan Qing

It is widely believed that engineering a model to be invariant/equivariant improves generalisation. Despite the growing popularity of this approach, a precise characterisation of the generalisation benefit is lacking. By considering the…

Machine Learning · Statistics 2021-07-07 Bryn Elesedy , Sheheryar Zaidi

Normalizing flows have shown great success as general-purpose density estimators. However, many real world applications require the use of domain-specific knowledge, which normalizing flows cannot readily incorporate. We propose…

Machine Learning · Statistics 2022-03-17 Gianluigi Silvestri , Emily Fertig , Dave Moore , Luca Ambrogioni

The learning properties of finite size polynomial Support Vector Machines are analyzed in the case of realizable classification tasks. The normalization of the high order features acts as a squeezing factor, introducing a strong anisotropy…

Disordered Systems and Neural Networks · Physics 2009-10-31 Sebastian Risau-Gusman , Mirta B. Gordon

Mixture-of-experts (MoE) model incorporates the power of multiple submodels via gating functions to achieve greater performance in numerous regression and classification applications. From a theoretical perspective, while there have been…

Machine Learning · Statistics 2024-06-25 Huy Nguyen , Pedram Akbarian , TrungTin Nguyen , Nhat Ho

Mixture of Experts (MoE) architectures have demonstrated remarkable success in scaling neural networks, yet their application to continual learning remains fundamentally limited by a critical vulnerability: the learned gating network itself…

Machine Learning · Computer Science 2025-12-15 Dev Vyas

To deal with changing environments, a new performance measure -- adaptive regret, defined as the maximum static regret over any interval, was proposed in online learning. Under the setting of online convex optimization, several algorithms…

Machine Learning · Computer Science 2025-08-04 Lijun Zhang , Wenhao Yang , Guanghui Wang , Wei Jiang , Zhi-Hua Zhou

We introduce a general framework for regression in the errors-in-variables regime, allowing for full flexibility about the dimensionality of the data, observational error probability density types, the (nonlinear) model type and the…

Methodology · Statistics 2024-11-19 Wolfgang Hoegele , Sarah Brockhaus

Predictive inference requires balancing statistical accuracy against informational complexity, yet the choice of complexity measure is usually imposed rather than derived. We treat econometric objects as predictive rules, mappings from…

Statistics Theory · Mathematics 2026-02-16 Nicholas G. Polson , Daniel Zantedeschi

E.T. Jaynes, originator of the maximum entropy interpretation of statistical mechanics, emphasized that there is an inevitable trade-off between the conflicting requirements of robustness and accuracy for any inferencing algorithm. This is…

Information Theory · Computer Science 2014-04-24 Kenric P. Nelson , Brian J. Scannell , Herbert Landau

We study the performance of general dynamic matching models. This model is defined by a connected graph, where nodes represent the class of items and the edges the compatibilities between items. Items of different classes arrive one by one…

Computer Science and Game Theory · Computer Science 2020-09-22 Arnaud Cadas , Josu Doncel , Jean-Michel Fourneau , Ana Bušić

A mixture preorder is a preorder on a mixture space (such as a convex set) that is compatible with the mixing operation. In decision theoretic terms, it satisfies the central expected utility axiom of strong independence. We consider when a…

Theoretical Economics · Economics 2021-02-16 David McCarthy , Kalle Mikkola , Teruji Thomas

We investigate the convergence properties of the EM algorithm when applied to overspecified Gaussian mixture models -- that is, when the number of components in the fitted model exceeds that of the true underlying distribution. Focusing on…

Machine Learning · Statistics 2025-06-16 Zhenisbek Assylbekov , Alan Legg , Artur Pak
‹ Prev 1 8 9 10 Next ›