English
Related papers

Related papers: Variable selection for Na\"ive Bayes classificatio…

200 papers

Gaussian empirical Bayes methods usually maintain a precision independence assumption: The unknown parameters of interest are independent from the known standard errors of the estimates. This assumption is often theoretically questionable…

Econometrics · Economics 2025-12-30 Jiafeng Chen

Sparse variational approximations are popular methods for scaling up inference and learning in Gaussian processes to larger datasets. For $N$ training points, exact inference has $O(N^3)$ cost; with $M \ll N$ features, state of the art…

Machine Learning · Statistics 2024-04-15 Talay M Cheema , Carl Edward Rasmussen

Automated feature selection is important for text categorization to reduce the feature size and to speed up the learning process of classifiers. In this paper, we present a novel and efficient feature selection framework based on the…

Machine Learning · Statistics 2016-11-15 Bo Tang , Steven Kay , Haibo He

Students opting for Engineering as their discipline is increasing rapidly. But due to various factors and inappropriate primary education in India, failure rates are high. Students are unable to excel in core engineering because of complex…

Machine Learning · Computer Science 2016-11-14 Muhammed Salman Shamsi , Jhansi Lakshmi

We propose a scalable Bayesian preference learning method for jointly predicting the preferences of individuals as well as the consensus of a crowd from pairwise labels. Peoples' opinions often differ greatly, making it difficult to predict…

Machine Learning · Computer Science 2019-12-13 Edwin Simpson , Iryna Gurevych

Feature selection is a crucial preprocessing step in data analytics and machine learning. Classical feature selection algorithms select features based on the correlations between predictive features and the class variable and do not attempt…

Machine Learning · Computer Science 2019-11-19 Kui Yu , Xianjie Guo , Lin Liu , Jiuyong Li , Hao Wang , Zhaolong Ling , Xindong Wu

We develop a Bayesian framework for variable selection in linear regression with autocorrelated errors, accommodating lagged covariates and autoregressive structures. This setting occurs in time series applications where responses depend on…

Methodology · Statistics 2025-08-18 Alokesh Manna , Sujit K. Ghosh

Multivariate matched proportions (MMP) data appears in a variety of contexts including post-market surveillance of adverse events in pharmaceuticals, disease classification, and agreement between care providers. It consists of multiple sets…

Methodology · Statistics 2023-05-08 Mark J. Meyer , Haobo Cheng , Katherine Hobbs Knutson

Theoretical results show that Bayesian methods can achieve lower bounds on regret for online logistic regression. In practice, however, such techniques may not be feasible especially for very large feature sets. Various approximations that,…

Machine Learning · Computer Science 2021-01-29 Gil I. Shamir , Wojciech Szpankowski

We develop a model-based empirical Bayes approach to variable selection problems in which the number of predictors is very large, possibly much larger than the number of responses (the so-called 'large p, small n' problem). We consider the…

Methodology · Statistics 2015-10-14 Haim Y. Bar , James G. Booth , Martin T. Wells

Equivariant models leverage prior knowledge on symmetries to improve predictive performance, but misspecified architectural constraints can harm it instead. While work has explored learning or relaxing constraints, selecting among…

Machine Learning · Computer Science 2025-07-16 Putri A. van der Linden , Alexander Timans , Dharmesh Tailor , Erik J. Bekkers

Variable screening methods have been shown to be effective in dimension reduction under the ultra-high dimensional setting. Most existing screening methods are designed to rank the predictors according to their individual contributions to…

Methodology · Statistics 2022-02-08 Ye Tian , Yang Feng

Many recently developed Bayesian methods have focused on sparse signal detection. However, much less work has been done addressing the natural follow-up question: how to make valid inferences for the magnitude of those signals after…

Methodology · Statistics 2021-03-02 Spencer Woody , Oscar Hernan Madrid Padilla , James G. Scott

Inspired by applications in sports where the skill of players or teams competing against each other varies over time, we propose a probabilistic model of pairwise-comparison outcomes that can capture a wide range of time dynamics. We…

Machine Learning · Statistics 2019-05-20 Lucas Maystre , Victor Kristof , Matthias Grossglauser

We propose a novel classification technique whose aim is to select an appropriate representation for each datapoint, in contrast to the usual approach of selecting a representation encompassing the whole dataset. This datum-wise…

Artificial Intelligence · Computer Science 2012-03-02 Gabriel Dulac-Arnold , Ludovic Denoyer , Philippe Preux , Patrick Gallinari

Variable selection has played a critical role in modern statistical learning and scientific discoveries. Numerous regularization and Bayesian variable selection methods have been developed in the past two decades for variable selection, but…

Methodology · Statistics 2024-03-04 Travis Canida , Hongjie Ke , Shuo Chen , Zhenayo Ye , Tianzhou Ma

The pathwise coordinate optimization is one of the most important computational frameworks for high dimensional convex and nonconvex sparse learning problems. It differs from the classical coordinate optimization algorithms in three salient…

Machine Learning · Statistics 2017-06-06 Tuo Zhao , Han Liu , Tong Zhang

Sparse linear discriminant analysis via penalized optimal scoring is a successful tool for classification in high-dimensional settings. While the variable selection consistency of sparse optimal scoring has been established, the…

Statistics Theory · Mathematics 2021-04-01 Irina Gaynanova

The discovery of partial differential equations (PDEs) is a challenging task that involves both theoretical and empirical methods. Machine learning approaches have been developed and used to solve this problem; however, it is important to…

Machine Learning · Statistics 2023-06-09 Kalpesh More , Tapas Tripura , Rajdip Nayek , Souvik Chakraborty

Recent work has focused on the problem of conducting linear regression when the number of covariates is very large, potentially greater than the sample size. To facilitate this, one useful tool is to assume that the model can be well…

Methodology · Statistics 2011-11-21 Zhou Fang