中文
相关论文

相关论文: Nonparametric Bayesian Knockoff Generators for Fea…

200 篇论文

Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery…

机器学习 · 统计学 2025-07-18 Omar Melikechi , David B. Dunson , Jeffrey W. Miller

In this paper, we study the problem of learning multi-dimensional Gaussian Mixture Models (GMMs), with a specific focus on model order selection and efficient mixing distribution estimation. We first establish an information-theoretic lower…

机器学习 · 统计学 2026-03-23 Xinyu Liu , Hai Zhang

Genome-wide association studies (GWAS) often find association signals between many genetic variants and traits of interest in a genomic region. Functional annotations of these variants provide valuable prior information that helps…

统计方法学 · 统计学 2026-01-07 Xiangyu Zhang , Lijun Wang , Changjun Li , Chen Lin , Hongyu Zhao

Predictions of nuclear properties far from measured data are inherently imprecise because of uncertainties in our knowledge of nuclear forces and in our treatment of quantum many-body effects in strongly-interacting systems. While the model…

核理论 · 物理学 2022-09-14 Rodrigo Navarro Perez , Nicolas Schunck

Algorithms that ensure reproducible findings from large-scale, high-dimensional data are pivotal in numerous signal processing applications. In recent years, multivariate false discovery rate (FDR) controlling methods have emerged,…

统计方法学 · 统计学 2024-01-31 Jasin Machkour , Michael Muma , Daniel P. Palomar

Kernel methods form a powerful, versatile, and theoretically-grounded unifying framework to solve nonlinear problems in signal processing and machine learning. The standard approach relies on the kernel trick to perform pairwise evaluations…

机器学习 · 计算机科学 2019-12-11 Kan Li , Jose C. Principe

Purpose: We propose a general framework for quantifying predictive uncertainties of dose-related quantities and leveraging this information in a dose mimicking problem in the context of automated radiation therapy treatment planning.…

医学物理 · 物理学 2021-09-08 Tianfang Zhang , Rasmus Bokrantz , Jimmy Olsson

This paper proposes a reliable neural network pruning algorithm by setting up a scientific control. Existing pruning methods have developed various hypotheses to approximate the importance of filters to the network and then execute filter…

计算机视觉与模式识别 · 计算机科学 2021-01-12 Yehui Tang , Yunhe Wang , Yixing Xu , Dacheng Tao , Chunjing Xu , Chao Xu , Chang Xu

We apply the knockoff procedure to factor selection in finance. By building fake but realistic factors, this procedure makes it possible to control the fraction of false discovery in a given set of factors. To show its versatility, we apply…

统计金融 · 定量金融 2021-07-07 Damien Challet , Christian Bongiorno , Guillaume Pelletier

Non-negative matrix factorization (NMF) is widely used in many applications for dimensionality reduction. Inferring an appropriate number of factors for NMF is a challenging problem, and several approaches based on information criteria or…

统计方法学 · 统计学 2025-02-18 Alessandro Zito , Jeffrey W. Miller

In many multiple testing applications in genetics, the signs of test statistics provide useful directional information, such as whether genes are potentially up- or down-regulated between two experimental conditions. However, most existing…

统计方法学 · 统计学 2025-07-22 Zhaoyang Tian , Kun Liang , Pengfei Li

Bayesian nonparametric methods are a popular choice for analysing survival data due to their ability to flexibly model the distribution of survival times. These methods typically employ a nonparametric prior on the survival function that is…

统计方法学 · 统计学 2022-02-22 Edwin Fong , Brieuc Lehmann

A fast forward feature selection algorithm is presented in this paper. It is based on a Gaussian mixture model (GMM) classifier. GMM are used for classifying hyperspectral images. The algorithm selects iteratively spectral features that…

计算机视觉与模式识别 · 计算机科学 2015-01-06 Mathieu Fauvel , Clement Dechesne , Anthony Zullo , Frédéric Ferraty

Simultaneously finding multiple influential variables and controlling the false discovery rate (FDR) for linear regression models is a fundamental problem. We here propose the Gaussian Mirror (GM) method, which creates for each predictor…

统计方法学 · 统计学 2021-03-22 Xin Xing , Zhigen Zhao , Jun S. Liu

In many application areas, data are collected on a categorical response and high-dimensional categorical predictors, with the goals being to build a parsimonious model for classification while doing inferences on the important predictors.…

统计方法学 · 统计学 2013-01-22 Yun Yang , David B. Dunson

Large-scale assessment data typically include numerous categorical variables, often affected by missing values. Motivated by the challenges arising in this framework, we extend the knockoffs method for selecting predictors to settings with…

统计方法学 · 统计学 2026-05-13 Silvia Bacci , Emanuela Dreassi , Leonardo Grilli , Carla Rampichini

Bayesian network classifiers provide a feasible solution to tabular data classification, with a number of merits like high time and memory efficiency, and great explainability. However, due to the parameter explosion and data sparsity…

机器学习 · 计算机科学 2025-08-18 Huan Zhang , Daokun Zhang , Kexin Meng , Geoffrey I. Webb

Although sparse autoencoders (SAEs) are crucial for identifying interpretable features in neural networks, it is still challenging to distinguish between real computational patterns and erroneous correlations. We introduce Model-X knockoffs…

机器学习 · 计算机科学 2025-11-18 Tsogt-Ochir Enkhbayar

Deep Gaussian processes (DGP) have appealing Bayesian properties, can handle variable-sized data, and learn deep features. Their limitation is that they do not scale well with the size of the data. Existing approaches address this using a…

机器学习 · 计算机科学 2019-05-20 Issam H. Laradji , Mark Schmidt , Vladimir Pavlovic , Minyoung Kim

High-dimensional variable selection, with many more covariates than observations, is widely documented in standard regression models, but there are still few tools to address it in non-linear mixed-effects models where data are collected…