中文
相关论文

相关论文: Generalised Mutual Information for Discriminative …

200 篇论文

Mutual information (MI) is a fundamental measure of statistical dependence, with a myriad of applications to information theory, statistics, and machine learning. While it possesses many desirable structural properties, the estimation of…

信息论 · 计算机科学 2021-10-19 Ziv Goldfeld , Kristjan Greenewald

Sliced Mutual Information (SMI) is widely used as a scalable alternative to mutual information for measuring non-linear statistical dependence. Despite its advantages, such as faster convergence, robustness to high dimensionality, and…

机器学习 · 计算机科学 2025-12-10 Alexander Semenenko , Ivan Butakov , Alexey Frolov , Ivan Oseledets

The aim of this work is to provide bounds connecting two probability measures of the same event using R\'enyi $\alpha$-Divergences and Sibson's $\alpha$-Mutual Information, a generalization of respectively the Kullback-Leibler Divergence…

信息论 · 计算机科学 2020-01-20 Amedeo Roberto Esposito , Michael Gastpar , Ibrahim Issa

There is no, nor will there ever be, single best clustering algorithm. Nevertheless, we would still like to be able to distinguish between methods that work well on certain task types and those that systematically underperform. Clustering…

机器学习 · 计算机科学 2025-10-16 Marek Gagolewski

Eliciting labels from crowds is a potential way to obtain large labeled data. Despite a variety of methods developed for learning from crowds, a key challenge remains unsolved: \emph{learning from crowds without knowing the information…

机器学习 · 计算机科学 2019-06-04 Peng Cao , Yilun Xu , Yuqing Kong , Yizhou Wang

Variational inference with a factorized Gaussian posterior estimate is a widely used approach for learning parameters and hidden variables. Empirically, a regularizing effect can be observed that is poorly understood. In this work, we show…

机器学习 · 计算机科学 2019-09-04 Julius Kunze , Louis Kirsch , Hippolyt Ritter , David Barber

Clustering categorical distributions in the finite-dimensional probability simplex is a fundamental task met in many applications dealing with normalized histograms. Traditionally, the differential-geometric structures of the probability…

机器学习 · 计算机科学 2021-11-22 Frank Nielsen , Ke Sun

Sequence-to-sequence neural network models for generation of conversational responses tend to generate safe, commonplace responses (e.g., "I don't know") regardless of the input. We suggest that the traditional objective function, i.e., the…

计算与语言 · 计算机科学 2016-06-14 Jiwei Li , Michel Galley , Chris Brockett , Jianfeng Gao , Bill Dolan

Sliced mutual information (SMI) is defined as an average of mutual information (MI) terms between one-dimensional random projections of the random variables. It serves as a surrogate measure of dependence to classic MI that preserves many…

信息论 · 计算机科学 2022-10-18 Ziv Goldfeld , Kristjan Greenewald , Theshani Nuradha , Galen Reeves

Information-maximization clustering learns a probabilistic classifier in an unsupervised manner so that mutual information between feature vectors and cluster assignments is maximized. A notable advantage of this approach is that it only…

机器学习 · 统计学 2011-12-06 Masashi Sugiyama , Makoto Yamada , Manabu Kimura , Hirotaka Hachiya

We study federated clustering, where interconnected devices collaboratively cluster the data points of private local datasets. Focusing on hard clustering via the k-means principle, we formulate federated k-means as an instance of…

机器学习 · 计算机科学 2026-01-29 Xu Yang , Salvatore Rastelli , Alexander Jung

In recent years, information-theoretic generalization bounds have gained increasing attention for analyzing the generalization capabilities of meta-learning algorithms. However, existing results are confined to two-step bounds, failing to…

机器学习 · 统计学 2025-10-14 Wen Wen , Tieliang Gong , Yuxin Dong , Zeyu Gao , Yong-Jin Liu

Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distillation setting, a teacher model provides soft predictions to guide the training of a…

信息论 · 计算机科学 2026-05-18 Bingying Li , Haiyun He

Distribution learning focuses on learning the probability density function from a set of data samples. In contrast, clustering aims to group similar objects together in an unsupervised manner. Usually, these two tasks are considered…

机器学习 · 计算机科学 2023-08-31 Guanfang Dong , Chenqiu Zhao , Anup Basu

The Mutual Information (MI) is an often used measure of dependency between two random variables utilized in information theory, statistics and machine learning. Recently several MI estimators have been proposed that can achieve parametric…

信息论 · 计算机科学 2018-11-26 Morteza Noshad , Yu Zeng , Alfred O. Hero

Information-theoretic generalization bounds based on the supersample construction are a central tool for algorithm-dependent generalization analysis in the batch i.i.d.~setting. However, existing supersample conditional mutual information…

机器学习 · 统计学 2026-05-13 Futoshi Futami , Masahiro Fujisawa

We are assisting at a growing interest in the development of learning architectures with application to digital communication systems. Herein, we consider the detection/decoding problem. We aim at developing an optimal neural architecture…

信息论 · 计算机科学 2022-09-02 Andrea M. Tonello , Nunzio A. Letizia

Loss-based clustering methods, such as k-means and its variants, are standard tools for finding groups in data. However, the lack of quantification of uncertainty in the estimated clusters is a disadvantage. Model-based clustering based on…

统计方法学 · 统计学 2020-06-11 Tommaso Rigon , Amy H. Herring , David B. Dunson

With increasing volume of data being used across machine learning tasks, the capability to target specific subsets of data becomes more important. To aid in this capability, the recently proposed Submodular Mutual Information (SMI) has been…

机器学习 · 计算机科学 2024-10-28 Nathan Beck , Truong Pham , Rishabh Iyer

Mutual Information (MI) is an useful tool for the recognition of mutual dependence berween data sets. Differen methods for the estimation of MI have been developed when both data sets are discrete or when both data sets are continuous. The…

应用统计 · 统计学 2017-08-30 Miguel A. Ré , Guillermo G. Aguirre Varela