中文
相关论文

相关论文: A Powerful and Precise Feature-level Filter using …

200 篇论文

Controlling the false discovery rate (FDR) is a popular approach to multiple testing, variable selection, and related problems of simultaneous inference. In many contemporary applications, models are not specified by discrete variables,…

统计理论 · 数学 2024-04-16 Mateo Díaz , Venkat Chandrasekaran

With the growing adoption of deep learning models in different real-world domains, including computational biology, it is often necessary to understand which data features are essential for the model's decision. Despite extensive recent…

机器学习 · 计算机科学 2022-10-04 Prashnna K Gyawali , Xiaoxia Liu , James Zou , Zihuai He

Large-scale assessment data typically include numerous categorical variables, often affected by missing values. Motivated by the challenges arising in this framework, we extend the knockoffs method for selecting predictors to settings with…

统计方法学 · 统计学 2026-05-13 Silvia Bacci , Emanuela Dreassi , Leonardo Grilli , Carla Rampichini

Controlling the false discovery rate (FDR) in variable selection becomes challenging when predictors are correlated, as existing methods often exclude all members of correlated groups and consequently perform poorly for prediction. We…

统计方法学 · 统计学 2026-03-03 Sarah Organ , Toby Kenney , Hong Gu

The mitigation of false positives is an important issue when conducting multiple hypothesis testing. The most popular paradigm for false positives mitigation in high-dimensional applications is via the control of the false discovery rate…

统计方法学 · 统计学 2018-07-17 Hien D. Nguyen , Yohan Yee , Geoffrey J. McLachlan , Jason P. Lerch

In many practical applications of multiple hypothesis testing using the False Discovery Rate (FDR), the given hypotheses can be naturally partitioned into groups, and one may not only want to control the number of false discoveries (wrongly…

统计方法学 · 统计学 2016-11-01 Rina Foygel Barber , Aaditya Ramdas

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

信号处理 · 电气工程与系统科学 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

User and item features of side information are crucial for accurate recommendation. However, the large number of feature dimensions, e.g., usually larger than 10^7, results in expensive storage and computational cost. This prohibits fast…

信息检索 · 计算机科学 2018-09-20 Han Liu , Xiangnan He , Fuli Feng , Liqiang Nie , Rui Liu , Hanwang Zhang

Feature selection (FS) is assumed to improve predictive performance and identify meaningful features in high-dimensional datasets. Surprisingly, small random subsets of features (0.02-1%) match or outperform the predictive performance of…

机器学习 · 计算机科学 2025-09-22 Bhavesh Neekhra , Debayan Gupta , Partha Pratim Chakrabarti

With the advancements in computing technology and web-based applications, data is increasingly generated in multi-dimensional form. This data is usually sparse due to the presence of a large number of users and fewer user interactions. To…

机器学习 · 计算机科学 2020-03-10 Thirunavukarasu Balasubramaniam , Richi Nayak , Chau Yuen

Controlling the false discovery rate (FDR) is a critical challenge in large-scale data analysis, particularly in the presence of outliers. A common practice involves imposing a Student-$t$ distribution to eliminate the influence of…

统计方法学 · 统计学 2026-04-01 Yoshiko Hayashi

The False Discovery Rate (FDR) is a new statistical procedure to control the number of mistakes made when performing multiple hypothesis tests, i.e. when comparing many data against a given model hypothesis. The key advantage of FDR is that…

Brain network topology, derived from functional magnetic resonance imaging (fMRI), holds promise for improving Alzheimer's disease (AD) diagnosis. Current methods primarily focus on lower-order topological features, often overlooking the…

几何拓扑 · 数学 2025-09-19 Dengyi Zhao , Shanyong Li , Yunping Wang , Chenfei Wang , Zhiheng Zhou , Guiying Yan , Xingqin Qi

Feature selection is crucial for pinpointing relevant features in high-dimensional datasets, mitigating the 'curse of dimensionality,' and enhancing machine learning performance. Traditional feature selection methods for classification use…

机器学习 · 计算机科学 2025-04-08 Rittwika Kansabanik , Adrian Barbu

Feature selection is a technique in statistical prediction modeling that identifies features in a record with a strong statistical connection to the target variable. Excluding features with a weak statistical connection to the target…

量子物理 · 物理学 2025-11-07 Andrew Vlasic , Hunter Grant , Salvatore Certo

Now that disease-modifying therapies for Alzheimer disease have been approved by regulatory agencies, the early, objective, and accurate clinical diagnosis of AD based on the lowest-cost measurement modalities possible has become an…

图像与视频处理 · 电气工程与系统科学 2025-11-13 Faisal Ahmed , Taymaz Akan , Fatih Gelir , Owen T. Carmichael , Elizabeth A. Disbrow , Steven A. Conrad , Mohammad A. N. Bhuiyan

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

统计方法学 · 统计学 2011-11-16 Jianqing Fan , Xu Han , Weijie Gu

Even when aggregate accuracy is high, state-of-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. Additional data collection may not help in addressing these…

计算与语言 · 计算机科学 2023-05-30 Zexue He , Marco Tulio Ribeiro , Fereshte Khani

The advent of artificial intelligence, especially the progress of deep neural networks, is expected to revolutionize genetic research and offer unprecedented potential to decode the complex relationships between genetic variants and disease…

定量方法 · 定量生物学 2023-12-13 Tingting Hou , Chang Jiang , Qing Lu

Due to the rapid innovation of technology and the desire to find and employ biomarkers for neurodegenerative disease, high-dimensional data classification problems are routinely encountered in neuroimaging studies. To avoid over-fitting and…

机器学习 · 统计学 2018-06-19 Shan Shi , Farouk Nathoo