中文
相关论文

相关论文: Nonparametric Variable Selection, Clustering and P…

200 篇论文

Variable selection is a procedure to attain the truly important predictors from inputs. Complex nonlinear dependencies and strong coupling pose great challenges for variable selection in high-dimensional data. In addition, real-world…

统计方法学 · 统计学 2023-07-04 Keyao Wang , Huiwen Wang , Jichang Zhao , Lihong Wang

There is a rich literature proposing methods and establishing asymptotic properties of Bayesian variable selection methods for parametric models, with a particular focus on the normal linear regression model and an increasing emphasis on…

统计理论 · 数学 2011-08-16 Suprateek Kundu , David B. Dunson

Gaussian process is a theoretically appealing model for nonparametric analysis, but its computational cumbersomeness hinders its use in large scale and the existing reduced-rank solutions are usually heuristic. In this work, we propose a…

机器学习 · 统计学 2015-11-25 Leo L. Duan , Xia Wang , Rhonda D. Szczesniak

Variable selection is central to high-dimensional data analysis, and various algorithms have been developed. Ideally, a variable selection algorithm shall be flexible, scalable, and with theoretical guarantee, yet most existing algorithms…

机器学习 · 统计学 2021-02-04 Xin He , Junhui Wang , Shaogao Lv

We propose the supervised hierarchical Dirichlet process (sHDP), a nonparametric generative model for the joint distribution of a group of observations and a response variable directly associated with that whole group. We compare the sHDP…

机器学习 · 统计学 2014-12-18 Andrew M. Dai , Amos J. Storkey

A limitation of many clustering algorithms is the requirement to tune adjustable parameters for each application or even for each dataset. Some techniques require an \emph{a priori} estimate of the number of clusters while density-based…

统计方法学 · 统计学 2016-05-20 Jeremy F. Magland , Alex H. Barnett

There is a widespread need for statistical methods that can analyze high-dimensional datasets with- out imposing restrictive or opaque modeling assumptions. This paper describes a domain-general data analysis method called CrossCat.…

人工智能 · 计算机科学 2015-12-07 Vikash Mansinghka , Patrick Shafto , Eric Jonas , Cap Petschulat , Max Gasner , Joshua B. Tenenbaum

Additive nonparametric regression models provide an attractive tool for variable selection in high dimensions when the relationship between the response and predictors is complex. They offer greater flexibility compared to parametric…

机器学习 · 统计学 2016-07-12 Garret Vo , Debdeep Pati

Finite Gaussian mixture models are widely used for model-based clustering of continuous data. Nevertheless, since the number of model parameters scales quadratically with the number of variables, these models can be easily…

统计方法学 · 统计学 2018-09-25 Michael Fop , Thomas Brendan Murphy , Luca Scrucca

Semi-supervised clustering is the task of clustering data points into clusters where only a fraction of the points are labelled. The true number of clusters in the data is often unknown and most models require this parameter as an input.…

机器学习 · 计算机科学 2013-09-27 Amar Shah , Zoubin Ghahramani

Gaussian processes have become a popular tool for nonparametric regression because of their flexibility and uncertainty quantification. However, they often use stationary kernels, which limit the expressiveness of the model and may be…

机器学习 · 计算机科学 2025-07-17 Zachary James , Joseph Guinness

We propose a nonparametric factorization approach for sparsely observed tensors. The sparsity does not mean zero-valued entries are massive or dominated. Rather, it implies the observed entries are very few, and even fewer with the growth…

机器学习 · 统计学 2021-11-04 Conor Tillinghast , Zheng Wang , Shandian Zhe

In this paper, we address the problem of conducting statistical inference in settings involving large-scale data that may be high-dimensional and contaminated by outliers. The high volume and dimensionality of the data require distributed…

机器学习 · 统计学 2022-11-30 Emadaldin Mozafari-Majd , Visa Koivunen

When fitting statistical models, some predictors are often found to be correlated with each other, and functioning together. Many group variable selection methods are developed to select the groups of predictors that are closely related to…

统计方法学 · 统计学 2021-03-25 Zhiyuan Li

The nonparametric formulation of density-based clustering, known as modal clustering, draws a correspondence between groups and the attraction domains of the modes of the density function underlying the data. Its probabilistic foundation…

统计方法学 · 统计学 2020-10-27 Federico Ferraccioli , Giovanna Menardi

We consider the estimation of Dirichlet Process Mixture Models (DPMMs) in distributed environments, where data are distributed across multiple computing nodes. A key advantage of Bayesian nonparametric models such as DPMMs is that they…

机器学习 · 统计学 2017-09-20 Ruohui Wang , Dahua Lin

We present a fast method for generating random samples according to a variable density Poisson-disc distribution. A minimum threshold distance is used to create a background grid array for keeping track of those points that might affect any…

图像与视频处理 · 电气工程与系统科学 2021-06-17 Nicholas Dwork , Corey A. Baron , Ethan M. I. Johnson , Daniel O'Connor , John M. Pauly , Peder E. Z. Larson

We consider the problem of sparse variable selection on high dimension heterogeneous data sets, which has been taking on renewed interest recently due to the growth of biological and medical data sets with complex, non-i.i.d. structures and…

统计方法学 · 统计学 2024-04-22 Hui Liu , Xiang Liu , Jing Diao , Wenting Ye , Xueling Liu , Dehui Wei

We develop nonparametric Bayesian modelling approaches for Poisson processes, using weighted combinations of structured beta densities to represent the point process intensity function. For a regular spatial domain, such as the unit square,…

统计方法学 · 统计学 2021-06-10 Chunyi Zhao , Athanasios Kottas

Dirichlet process mixtures are flexible non-parametric models, particularly suited to density estimation and probabilistic clustering. In this work we study the posterior distribution induced by Dirichlet process mixtures as the sample size…

统计理论 · 数学 2022-11-29 Filippo Ascolani , Antonio Lijoi , Giovanni Rebaudo , Giacomo Zanella