中文
相关论文

相关论文: An Information-theoretic Approach to Unsupervised …

200 篇论文

The commonly used latent space embedding techniques, such as Principal Component Analysis, Factor Analysis, and manifold learning techniques, are typically used for learning effective representations of homogeneous data. However, they do…

机器学习 · 计算机科学 2021-10-04 Yasin Yilmaz , Mehmet Aktukmak , Alfred O. Hero

Integrative learning of multiple datasets has the potential to mitigate the challenge of small $n$ and large $p$ that is often encountered in analysis of big biomedical data such as genomics data. Detection of weak yet important signals can…

统计方法学 · 统计学 2022-07-04 Changgee Chang , Zongyu Dai , Jihwan Oh , Qi Long

A novel method for common and individual feature analysis from exceedingly large-scale data is proposed, in order to ensure the tractability of both the computation and storage and thus mitigate the curse of dimensionality, a major…

信号处理 · 电气工程与系统科学 2017-11-03 Ilia Kisil , Giuseppe G. Calvi , Danilo P. Mandic

Dynamic feature selection, where we sequentially query features to make accurate predictions with a minimal budget, is a promising paradigm to reduce feature acquisition costs and provide transparency into a model's predictions. The problem…

机器学习 · 计算机科学 2024-09-10 Soham Gadgil , Ian Covert , Su-In Lee

The ability to find short representations, i.e. to compress data, is crucial for many intelligent systems. We present a theory of incremental compression showing that arbitrary data strings, that can be described by a set of features, can…

信息论 · 计算机科学 2020-09-15 Arthur Franz , Oleksandr Antonenko , Roman Soletskyi

We propose learning discrete structured representations from unlabeled data by maximizing the mutual information between a structured latent variable and a target variable. Calculating mutual information is intractable in this setting. Our…

机器学习 · 计算机科学 2020-07-17 Karl Stratos , Sam Wiseman

Extraction of structure, in particular of group symmetries, is increasingly crucial to understanding and building intelligent models. In particular, some information-theoretic models of parsimonious learning have been argued to induce…

信息论 · 计算机科学 2025-07-08 Hippolyte Charvin , Nicola Catenacci Volpi , Daniel Polani

A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be…

机器学习 · 计算机科学 2019-09-12 Jonas Mueller , Alex Smola

The goal of data selection is to capture the most structural information from a set of data. This paper presents a fast and accurate data selection method, in which the selected samples are optimized to span the subspace of all data. We…

计算机视觉与模式识别 · 计算机科学 2018-11-30 Mohsen Joneidi , Alireza Zaeemzadeh , Nazanin Rahnavard , Mubarak Shah

How to extract useful insights from data is always a challenge, especially if the data is multidimensional. Often, the data can be organized according to certain hierarchical structure that are stemmed either from data collection process or…

应用统计 · 统计学 2016-04-21 Kun Yang , Wing Hung Wong

We are interested in learning data-driven representations that can generalize well, even when trained on inherently biased data. In particular, we face the case where some attributes (bias) of the data, if learned by the model, can severely…

计算机视觉与模式识别 · 计算机科学 2020-03-17 Ruggero Ragonesi , Riccardo Volpi , Jacopo Cavazza , Vittorio Murino

We investigate the problem of selecting features for datasets that can be naturally partitioned into subgroups (e.g., according to socio-demographic groups and age), each with its own dominant set of features. Within this subgroup-oriented…

机器学习 · 计算机科学 2024-12-10 Bar Genossar , Thinh On , Md. Mouinul Islam , Ben Eliav , Senjuti Basu Roy , Avigdor Gal

Measuring the relationship between any pair of variables is a rich and active area of research that is central to scientific practice. In contrast, characterizing the common information among any group of variables is typically a…

机器学习 · 统计学 2017-06-20 Greg Ver Steeg , Shuyang Gao , Kyle Reing , Aram Galstyan

In many complex applications, data heterogeneity and homogeneity exist simultaneously. Ignoring either one will result in incorrect statistical inference. In addition, coping with complex data that are non-Euclidean becomes more common. To…

统计方法学 · 统计学 2021-05-28 Zixuan Han , Tao Li , Jinhong You

One of the goals of probabilistic inference is to decide whether an empirically observed distribution is compatible with a candidate Bayesian network. However, Bayesian networks with hidden variables give rise to highly non-trivial…

机器学习 · 统计学 2014-10-14 R. Chaves , L. Luft , T. O. Maciel , D. Gross , D. Janzing , B. Schölkopf

We propose a method to facilitate exploration and analysis of new large data sets. In particular, we give an unsupervised deep learning approach to learning a latent representation that captures semantic similarity in the data set. The core…

计算机视觉与模式识别 · 计算机科学 2020-12-23 Gary B Huang , Huei-Fang Yang , Shin-ya Takemura , Pat Rivlin , Stephen M Plaza

Estimating treatment effects from observational data is challenging due to two main reasons: (a) hidden confounding, and (b) covariate mismatch (control and treatment groups not having identical distributions). Long lines of works exist…

机器学习 · 计算机科学 2025-04-30 Praharsh Nanavati , Ranjitha Prasad , Karthikeyan Shanmugam

From a machine learning point of view, identifying a subset of relevant features from a real data set can be useful to improve the results achieved by classification methods and to reduce their time and space complexity. To achieve this…

机器学习 · 计算机科学 2017-05-23 Pietro Cassara , Alessandro Rozza , Mirco Nanni

We introduce instancewise feature selection as a methodology for model interpretation. Our method is based on learning a function to extract a subset of features that are most informative for each given example. This feature selector is…

机器学习 · 计算机科学 2018-06-15 Jianbo Chen , Le Song , Martin J. Wainwright , Michael I. Jordan

We formulate meta learning using information theoretic concepts; namely, mutual information and the information bottleneck. The idea is to learn a stochastic representation or encoding of the task description, given by a training set, that…