中文
相关论文

相关论文: An Information-theoretic Approach to Unsupervised …

200 篇论文

Feature selection by maximizing high-order mutual information between the selected feature vector and a target variable is the gold standard in terms of selecting the best subset of relevant features that maximizes the performance of…

机器学习 · 计算机科学 2022-10-19 Magda Amiridi , Nikos Kargas , Nicholas D. Sidiropoulos

Estimating how a treatment affects units individually, known as heterogeneous treatment effect (HTE) estimation, is an essential part of decision-making and policy implementation. The accumulation of large amounts of data in many domains,…

机器学习 · 计算机科学 2022-06-28 Christopher Tran , Elena Zheleva

Algorithmic fairness is becoming increasingly important in data mining and machine learning. Among others, a foundational notation is group fairness. The vast majority of the existing works on group fairness, with a few exceptions,…

机器学习 · 计算机科学 2023-01-03 Jian Kang , Tiankai Xie , Xintao Wu , Ross Maciejewski , Hanghang Tong

A wide range of systems exhibit high dimensional incomplete data. Accurate estimation of the missing data is often desired, and is crucial for many downstream analyses. Many state-of-the-art recovery methods involve supervised learning…

计算机视觉与模式识别 · 计算机科学 2019-03-15 Adrian V. Dalca , John Guttag , Mert R. Sabuncu

A fundamental task in science is to determine the underlying causal relations because it is the knowledge of this functional structure what leads to the correct interpretation of an effect given the apparent associations in the observed…

人工智能 · 计算机科学 2024-08-02 Alexandre Trilla , Nenad Mijatovic

Complex categorical data is often hierarchically coupled with heterogeneous relationships between attributes and attribute values and the couplings between objects. Such value-to-object couplings are heterogeneous with complementary and…

机器学习 · 计算机科学 2020-07-28 Chengzhang Zhu , Longbing Cao , Jianping Yin

Functional principal component analysis (FPCA) is a fundamental tool and has attracted increasing attention in recent decades, while existing methods are restricted to data with a single or finite number of random functions (much smaller…

统计方法学 · 统计学 2021-01-22 Xiaoyu Hu , Fang Yao

Distributed learning of probabilistic models from multiple data repositories with minimum communication is increasingly important. We study a simple communication-efficient learning framework that first calculates the local maximum…

机器学习 · 统计学 2014-10-13 Qiang Liu , Alexander Ihler

We consider a set of probabilistic functions of some input variables as a representation of the inputs. We present bounds on how informative a representation is about input data. We extend these bounds to hierarchical representations so…

机器学习 · 统计学 2015-02-03 Greg Ver Steeg , Aram Galstyan

Estimating heterogeneous treatment effects is an important problem across many domains. In order to accurately estimate such treatment effects, one typically relies on data from observational studies or randomized experiments. Currently,…

When a machine-learning algorithm makes biased decisions, it can be helpful to understand the sources of disparity to explain why the bias exists. Towards this, we examine the problem of quantifying the contribution of each individual…

机器学习 · 计算机科学 2022-06-20 Sanghamitra Dutta , Praveen Venkatesh , Pulkit Grover

We investigate the task of retrieving information from compositional distributed representations formed by Hyperdimensional Computing/Vector Symbolic Architectures and present novel techniques which achieve new information rate bounds.…

Form a pure mathematical point of view, common functional forms representing different physical phenomena can be defined. For example, rates of chemical reactions, diffusion and heat transfer are all governed by exponential-type…

机器学习 · 计算机科学 2019-10-01 Navid Zobeiry , Keith D. Humfeld

Information theory is an outstanding framework to measure uncertainty, dependence and relevance in data and systems. It has several desirable properties for real world applications: it naturally deals with multivariate data, it can handle…

This paper introduces a general Bayesian non- parametric latent feature model suitable to per- form automatic exploratory analysis of heterogeneous datasets, where the attributes describing each object can be either discrete, continuous or…

机器学习 · 统计学 2017-07-27 Isabel Valera , Melanie F. Pradier , Zoubin Ghahramani

Hidden variable graphical models can sometimes imply constraints on the observable distribution that are more complex than simple conditional independence relations. These observable constraints can falsify assumptions of the model that…

统计方法学 · 统计学 2026-05-12 Michael C. Sachs , Erin E. Gabriel , Robin J. Evans , Arvid Sjölander

Understanding the contribution of individual features in predictive models remains a central goal in interpretable machine learning, and while many model-agnostic methods exist to estimate feature importance, they often fall short in…

机器学习 · 计算机科学 2025-07-08 Ivan Lazic , Chiara Barà , Marta Iovino , Sebastiano Stramaglia , Niksa Jakovljevic , Luca Faes

We introduce an information-theoretic framework that views learning as universal prediction under log loss, characterized through regret bounds. Central to the framework is an effective notion of architecture-based model complexity, defined…

机器学习 · 计算机科学 2025-11-04 Meir Feder , Ruediger Urbanke , Yaniv Fogel

Learning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models…

机器学习 · 计算机科学 2020-10-05 Ruixiang Zhang , Masanori Koyama , Katsuhiko Ishiguro

Analysis of a probabilistic system often requires to learn the joint probability distribution of its random variables. The computation of the exact distribution is usually an exhaustive precise analysis on all executions of the system. To…

信息论 · 计算机科学 2023-07-19 Fabrizio Biondi , Yusuke Kawamoto , Axel Legay , Louis-Marie Traonouez