中文
相关论文

相关论文: Feature Selection via the Intervened Interpolative…

200 篇论文

Variable selection is crucial for sparse modeling in this age of big data. Missing values are common in data, and make variable selection more complicated. The approach of multiple imputation (MI) results in multiply imputed datasets for…

统计方法学 · 统计学 2025-09-04 Yong-Shiuan Lee

In this paper we propose a bayesian approach for near-duplicate image detection, and investigate how different probabilistic models affect the performance obtained. The task of identifying an image whose metadata are missing is often…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Lucas Moutinho Bueno , Eduardo Valle , Ricardo da Silva Torres

This paper presents a randomized algorithm for computing the near-optimal low-rank dynamic mode decomposition (DMD). Randomized algorithms are emerging techniques to compute low-rank matrix approximations at a fraction of the cost of…

数值分析 · 数学 2019-11-28 N. Benjamin Erichson , Lionel Mathelin , Steven L. Brunton , J. Nathan Kutz

Medical imaging involves high-dimensional data, yet their acquisition is obtained for limited samples. Multivariate predictive models have become popular in the last decades to fit some external variables from imaging data, and standard…

应用统计 · 统计学 2018-06-18 Jérôme-Alexis Chevalier , Joseph Salmon , Bertrand Thirion

We consider the problem of learning from data corrupted by underrepresentation bias, where positive examples are filtered from the data at different, unknown rates for a fixed number of sensitive groups. We show that with a small amount of…

机器学习 · 计算机科学 2024-06-05 Emily Diana , Alexander Williams Tolbert

This paper introduces a "kernel-independent" interpolative decomposition butterfly factorization (IDBF) as a data-sparse approximation for matrices that satisfy a complementary low-rank property. The IDBF can be constructed in $O(N\log N)$…

数值分析 · 数学 2018-10-09 Qiyuan Pang , Kenneth L. Ho , Haizhao Yang

Incorporating feature selection into a classification or regression method often carries a number of advantages. In this paper we formalize feature selection specifically from a discriminative perspective of improving…

机器学习 · 计算机科学 2013-01-18 Tony S. Jebara , Tommi S. Jaakkola

Variable selection in high-dimensional space characterizes many contemporary problems in scientific discovery and decision making. Many frequently-used techniques are based on independence screening; examples include correlation ranking…

统计方法学 · 统计学 2008-12-18 Jianqing Fan , Richard Samworth , Yichao Wu

Inpainting-based image compression is a promising alternative to classical transform-based lossy codecs. Typically it stores a carefully selected subset of all pixel locations and their colour values. In the decoding phase the missing…

图像与视频处理 · 电气工程与系统科学 2023-05-16 Ferdinand Jost , Vassillen Chizhov , Joachim Weickert

Tucker tensor decomposition offers a more effective representation for multiway data compared to the widely used PARAFAC model. However, its flexibility brings the challenge of selecting the appropriate latent multi-rank. To overcome the…

统计方法学 · 统计学 2025-05-19 Federica Stolf , Antonio Canale

An interpolation-based decoding scheme for interleaved subspace codes is presented. The scheme can be used as a (not necessarily polynomial-time) list decoder as well as a probabilistic unique decoder. Both interpretations allow to decode…

信息论 · 计算机科学 2014-08-07 Hannes Bartz , Antonia Wachter-Zeh

In computational biology, gene expression datasets are characterized by very few individual samples compared to a large number of measurements per sample. Thus, it is appealing to merge these datasets in order to increase the number of…

统计方法学 · 统计学 2011-08-18 Meili Baragatti

Compression is a crucial solution for data reduction in modern scientific applications due to the exponential growth of data from simulations, experiments, and observations. Compression with progressive retrieval capability allows users to…

分布式、并行与集群计算 · 计算机科学 2025-04-08 Zhuoxun Yang , Sheng Di , Longtao Zhang , Ruoyu Li , Ximiao Li , Jiajun Huang , Jinyang Liu , Franck Cappello , Kai Zhao

We introduce a novel Bayesian hybrid matrix factorisation model (HMF) for data integration, based on combining multiple matrix factorisation methods, that can be used for in- and out-of-matrix prediction of missing values. The model is very…

机器学习 · 统计学 2017-04-18 Thomas Brouwer , Pietro Lió

Information theoretic active learning has been widely studied for probabilistic models. For simple regression an optimal myopic policy is easily tractable. However, for other tasks and with more complex models, such as classification with…

机器学习 · 统计学 2011-12-30 Neil Houlsby , Ferenc Huszár , Zoubin Ghahramani , Máté Lengyel

Intrusion detection system (IDS) is one of extensively used techniques in a network topology to safeguard the integrity and availability of sensitive assets in the protected systems. Although many supervised and unsupervised learning…

密码学与安全 · 计算机科学 2020-04-03 Yuyang Zhou , Guang Cheng , Shanqing Jiang , Mian Dai

When a machine-learning algorithm makes biased decisions, it can be helpful to understand the sources of disparity to explain why the bias exists. Towards this, we examine the problem of quantifying the contribution of each individual…

机器学习 · 计算机科学 2022-06-20 Sanghamitra Dutta , Praveen Venkatesh , Pulkit Grover

Extracting demographic features from hidden factors is an innovative concept that provides multiple and relevant applications. The matrix factorization model generates factors which do not incorporate semantic knowledge. This paper provides…

信息检索 · 计算机科学 2020-12-22 Jesús Bobadilla , Ángel González-Prieto , Fernando Ortega , Raúl Lara-Cabrera

We study Bayesian inference methods for solving linear inverse problems, focusing on hierarchical formulations where the prior or the likelihood function depend on unspecified hyperparameters. In practice, these hyperparameters are often…

数值分析 · 数学 2018-08-01 Qingping Zhou , Wenqing Liu , Jinglai Li , Youssef M. Marzouk

We propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many…