English
Related papers

Related papers: A Powerful and Precise Feature-level Filter using …

200 papers

In many scientific settings there is a need for adaptive experimental design to guide the process of identifying regions of the search space that contain as many true positives as possible subject to a low rate of false discoveries (i.e.…

Machine Learning · Statistics 2020-08-18 Lalit Jain , Kevin Jamieson

The detection of small objects is a challenging task in computer vision. Conventional object detection methods have difficulty in finding the balance between high detection and low false alarm rates. In the literature, some methods have…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Alina Ciocarlan , Sylvie Le Hegarat-Mascle , Sidonie Lefebvre , Arnaud Woiselle

Replicability is a fundamental quality of scientific discoveries: we are interested in those signals that are detectable in different laboratories, study populations, across time etc. Unlike meta-analysis which accounts for experimental…

Methodology · Statistics 2021-11-19 Jingshu Wang , Lin Gui , Weijie J. Su , Chiara Sabatti , Art B. Owen

For precision medicine and personalized treatment, we need to identify predictive markers of disease. We focus on Alzheimer's disease (AD), where magnetic resonance imaging scans provide information about the disease status. By combining…

Machine Learning · Statistics 2019-03-06 Stefan Konigorski , Shahryar Khorasani , Christoph Lippert

Identifying which variables do influence a response while controlling false positives pervades statistics and data science. In this paper, we consider a scenario in which we only have access to summary statistics, such as the values of…

Alzheimer's patients gradually lose their ability to think, behave, and interact with others. Medical history, laboratory tests, daily activities, and personality changes can all be used to diagnose the disorder. A series of time-consuming…

Machine Learning · Computer Science 2022-12-02 Md. Sharifur Rahman , Professor Girijesh Prasad

Model-X knockoffs is a wrapper that transforms essentially any feature importance measure into a variable selection algorithm, which discovers true effects while rigorously controlling the expected fraction of false positives. A frequently…

Methodology · Statistics 2024-03-12 Stephen Bates , Emmanuel Candès , Lucas Janson , Wenshuo Wang

There is a challenge in selecting high-dimensional mediators when the mediators have complex correlation structures and interactions. In this work, we frame the high-dimensional mediator selection problem into a series of hypothesis tests…

Methodology · Statistics 2025-09-16 Runqiu Wang , Ran Dai , Jieqiong Wang , Kah Meng Soh , Ziyang Xu , Mohamed Azzam , Hongying Dai , Cheng Zheng

Selecting important features in high-dimensional survival analysis is critical for identifying confirmatory biomarkers while maintaining rigorous error control. In this paper, we propose a derandomized knockoffs procedure for Cox regression…

Methodology · Statistics 2025-12-15 Rui Liu , Nan Sun

Ensemble learning use multiple algorithms to obtain better predictive performance than any single one of its constituent algorithms could. With growing popularity of deep learning, researchers have started to ensemble them for various…

Machine Learning · Computer Science 2019-05-31 Ning An , Huitong Ding , Jiaoyun Yang , Rhoda Au , Ting Fang Alvin Ang

Replicability analysis aims to identify the findings that replicated across independent studies that examine the same features. We provide powerful novel replicability analysis procedures for two studies for FWER and for FDR control on the…

Methodology · Statistics 2019-03-01 Marina Bogomolov , Ruth Heller

Click-through rate (CTR) prediction plays important role in personalized advertising and recommender systems. Though many models have been proposed such as FM, FFM and DeepFM in recent years, feature engineering is still a very important…

Information Retrieval · Computer Science 2021-07-27 Qingyun She , Zhiqiang Wang , Junlin Zhang

Testing for differences in features between clusters in various applications often leads to inflated false positives when practitioners use the same dataset to identify clusters and then test features, an issue commonly known as ``double…

Methodology · Statistics 2024-10-10 Lijun Wang , Yingxin Lin , Hongyu Zhao

We apply the knockoff procedure to factor selection in finance. By building fake but realistic factors, this procedure makes it possible to control the fraction of false discovery in a given set of factors. To show its versatility, we apply…

Statistical Finance · Quantitative Finance 2021-07-07 Damien Challet , Christian Bongiorno , Guillaume Pelletier

Randomization tests are a popular method for testing causal effects in clinical trials with finite-sample validity. In the presence of heterogeneous treatment effects, it is often of interest to select a subgroup that benefits from the…

Methodology · Statistics 2025-04-29 Zijun Gao

More than 10.7% of people aged 65 and older are affected by Alzheimer's disease. Early diagnosis and treatment are crucial as most Alzheimer's patients are unaware of having it until the effects become detrimental. AI has been known to use…

Image and Video Processing · Electrical Eng. & Systems 2023-09-19 Audrey Paleczny , Shubham Parab , Maxwell Zhang

Feature selection is a common step in many ranking, classification, or prediction tasks and serves many purposes. By removing redundant or noisy features, the accuracy of ranking or classification can be improved and the computational cost…

Information Retrieval · Computer Science 2022-05-10 Maurizio Ferrari Dacrema , Fabio Moroni , Riccardo Nembrini , Nicola Ferro , Guglielmo Faggioli , Paolo Cremonesi

Model-X knockoffs is a flexible wrapper method for high-dimensional regression algorithms, which provides guaranteed control of the false discovery rate (FDR). Due to the randomness inherent to the method, different runs of model-X…

Methodology · Statistics 2023-09-01 Zhimei Ren , Rina Foygel Barber

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

Methodology · Statistics 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

In artificial neural networks, understanding the contributions of input features on the prediction fosters model explainability and delivers relevant information about the dataset. While typical setups for feature importance ranking assess…