中文
相关论文

相关论文: Multilayer Knockoff Filter: Controlled variable se…

200 篇论文

Deep learning has become increasingly popular in both supervised and unsupervised machine learning thanks to its outstanding empirical performance. However, because of their intrinsic complexity, most deep learning methods are largely…

机器学习 · 计算机科学 2018-09-07 Yang Young Lu , Yingying Fan , Jinchi Lv , William Stafford Noble

Collaborative filtering is an effective recommendation approach in which the preference of a user on an item is predicted based on the preferences of other users with similar interests. A big challenge in using collaborative filtering…

信息检索 · 计算机科学 2012-03-19 Yu Zhang , Bin Cao , Dit-Yan Yeung

We provide new non-asymptotic false discovery proportion (FDP) confidence envelopes in several multiple testing settings relevant for modern high dimensional-data methods. We revisit the multiple testing scenarios considered in the recent…

统计理论 · 数学 2024-09-18 Iqraa Meah , Gilles Blanchard , Etienne Roquain

The genetic basis of multiple phenotypes such as gene expression, metabolite levels, or imaging features is often investigated by testing a large collection of hypotheses, probing the existence of association between each of the traits and…

应用统计 · 统计学 2015-04-06 Christine Peterson , Marina Bogomolov , Yoav Benjamini , Chiara Sabatti

False discovery rate (FDR) is a cornerstone of modern multiple testing. However, it often fails to guarantee the reliability of "marginal" discoveries that lie at the boundary of the rejection set, which are often crucial in high-precision…

统计方法学 · 统计学 2026-05-12 Yifan Zhang , Wentao Zhang , Changliang Zou , Haojie Ren

Majority voting is a simple mathematical function that returns the value that appears most often in a set. As a popular decision fusion technique, the majority voting function (MVF) finds applications in resolving conflicts, where a number…

机器学习 · 统计学 2023-09-19 Sina Aeeneh , Nikola Zlatanov , Jiangshan Yu

Improved procedures, in terms of smaller missed discovery rates (MDR), for performing multiple hypotheses testing with weak and strong control of the family-wise error rate (FWER) or the false discovery rate (FDR) are developed and studied.…

统计理论 · 数学 2011-03-10 Edsel A. Peña , Joshua D. Habiger , Wensong Wu

In this work, we present a new technique for the decomposition of multivariate data, which we call Multivariate Fast Iterative Filtering (MvFIF) algorithm. We study its properties, proving rigorously that it converges in finite time when…

数值分析 · 数学 2021-11-04 Antonio Cicone , Enza Pellegrino

The decomposition of a signal is a fundamental tool in many fields of research, including signal processing, geophysics, astrophysics, engineering, medicine, and many more. By breaking down complex signals into simpler oscillatory…

数值分析 · 数学 2024-12-03 Roberto Cavassi , Antonio Cicone , Enza Pellegrino , Haomin Zhou

The complexity of deep neural networks (DNNs) makes them powerful but also makes them challenging to interpret, hindering their applicability in error-intolerant domains. Existing methods attempt to reason about the internal mechanism of…

机器学习 · 计算机科学 2023-09-28 Winston Chen , William Stafford Noble , Yang Young Lu

Feature selection is essential for high-dimensional biomedical data, enabling stronger predictive performance, reduced computational cost, and improved interpretability in precision medicine applications. Existing approaches face notable…

机器学习 · 计算机科学 2026-01-07 Xiaoyan Sun , Qingyu Meng , Yalu Wen

The $\gamma$-FDP and $k$-FWER multiple testing error metrics, which are tail probabilities of the respective error statistics, have become popular recently as less-stringent alternatives to the FDR and FWER. We propose general and flexible…

统计方法学 · 统计学 2016-12-20 Jay Bartroff

Interpretable machine learning seeks to understand the reasoning process of complex black-box systems that are long notorious for lack of explainability. One flourishing approach is through counterfactual explanations, which provide…

人工智能 · 计算机科学 2023-06-02 Vy Vo , Trung Le , Van Nguyen , He Zhao , Edwin Bonilla , Gholamreza Haffari , Dinh Phung

We develop an extension of the Knockoff Inference procedure, introduced by Barber and Candes (2015). This new method, called Aggregation of Multiple Knockoffs (AKO), addresses the instability inherent to the random nature of Knockoff-based…

Molecular data from tumor profiles is high dimensional. Tumor profiles can be characterized by tens of thousands of gene expression features. Due to the size of the gene expression feature set machine learning methods are exposed to noisy…

机器学习 · 计算机科学 2020-07-14 Martin Palazzo , Pierre Beauseroy , Patricio Yankilevich

Knowing the features of a complex system that are highly relevant to a particular target variable is of fundamental interest in many areas of science. Existing approaches are often limited to linear settings, sometimes lack guarantees, and…

机器学习 · 计算机科学 2023-07-06 Francesco Quinzan , Ashkan Soleymani , Patrick Jaillet , Cristian R. Rojas , Stefan Bauer

Multi-view datasets offer diverse forms of data that can enhance prediction models by providing complementary information. However, the use of multi-view data leads to an increase in high-dimensional data, which poses significant challenges…

神经与进化计算 · 计算机科学 2024-03-05 Vandad Imani , Carlos Sevilla-Salcedo , Elaheh Moradi , Vittorio Fortino , Jussi Tohka

High-dimensional variable selection has emerged as one of the prevailing statistical challenges in the big data revolution. Many variable selection methods have been adapted for identifying single nucleotide polymorphisms (SNPs) linked to…

统计方法学 · 统计学 2024-08-21 Justin J. Van Ee , Diana Gamba , Jesse R. Lasky , Megan L. Vahsen , Mevin B. Hooten

Controlling the false discovery rate (FDR) in high-dimensional variable selection requires balancing rigorous error control with statistical power. Existing methods with provable guarantees are often overly conservative, creating a…

统计方法学 · 统计学 2026-02-06 Arnau Vilella , Jasin Machkour , Michael Muma , Daniel P. Palomar

Feature selection is critical in machine learning to reduce dimensionality and improve model accuracy and efficiency. The exponential growth in feature space dimensionality for modern datasets directly results in ambiguous samples and…

量子物理 · 物理学 2023-11-30 Haiyan Wang