English
Related papers

Related papers: One-at-a-time knockoffs: controlled false discover…

200 papers

Sorted L-One Penalized Estimation (SLOPE) has shown the nice theoretical property as well as empirical behavior recently on the false discovery rate (FDR) control of high-dimensional feature selection by adaptively imposing the…

Statistics Theory · Mathematics 2023-02-22 Jingxuan Liang , Hong Chen , Xuelin Zhang , Weifu Li , Xin Tang

We introduce Time-Conditioned Contraction Matching (TCCM), a novel method for semi-supervised anomaly detection in tabular data. TCCM is inspired by flow matching, a recent generative modeling framework that learns velocity fields between…

Machine Learning · Computer Science 2025-10-22 Zhong Li , Qi Huang , Yuxuan Zhu , Lincen Yang , Mohammad Mohammadi Amiri , Niki van Stein , Matthijs van Leeuwen

Principal component analysis (PCA) is one of the most powerful tools in machine learning. The simplest method for PCA, the power iteration, requires $\mathcal O(1/\Delta)$ full-data passes to recover the principal component of a matrix with…

Optimization and Control · Mathematics 2017-07-11 Christopher De Sa , Bryan He , Ioannis Mitliagkas , Christopher Ré , Peng Xu

Detecting anomalies in large, distributed systems presents several challenges. The first challenge arises from the sheer volume of data that needs to be processed. Flagging anomalies in a high-throughput environment calls for a careful…

Machine Learning · Computer Science 2025-10-07 Anupam Panwar , Himadri Pal , Jiali Chen , Kyle Cho , Riddick Jiang , Miao Zhao , Rajiv Krishnamurthy

We investigate the onset of chaos in a periodically kicked Dicke model (KDM), using the out-of-time-order correlator (OTOC) as a diagnostic tool, in both the oscillator and the spin subspaces. In the large spin limit, the classical…

Statistical Mechanics · Physics 2021-04-26 Sudip Sinha , Sayak Ray , Subhasis Sinha

This paper presents a survey on some recent advances for the type I error rate control in multiple testing methodology. We consider the problem of controlling the $k$-family-wise error rate (kFWER, probability to make $k$ false discoveries…

Methodology · Statistics 2011-03-15 Etienne Roquain

Many important tasks of large-scale recommender systems can be naturally cast as testing multiple linear forms for noisy matrix completion. These problems, however, present unique challenges because of the subtle bias-and-variance tradeoff…

Methodology · Statistics 2025-03-12 Wanteng Ma , Lilun Du , Dong Xia , Ming Yuan

Probing the out-of-equilibrium dynamics of quantum matter has gained renewed interest owing to immense experimental progress in artifcial quantum systems. Dynamical quantum measures such as the growth of entanglement entropy (EE) and…

Disordered Systems and Neural Networks · Physics 2018-04-04 Pranjal Bordia , Fabien Alet , Pavan Hosur

We analyze principal component regression (PCR) in a high-dimensional error-in-variables setting with fixed design. Under suitable conditions, we show that PCR consistently identifies the unique model with minimum $\ell_2$-norm. These…

Statistics Theory · Mathematics 2023-08-28 Anish Agarwal , Devavrat Shah , Dennis Shen

The complexity of modern electro-mechanical systems require the development of sophisticated diagnostic methods like anomaly detection capable of detecting deviations. Conventional anomaly detection approaches like signal processing and…

Machine Learning · Computer Science 2025-01-07 Abhishek Srinivasan , Varun Singapuri Ravi , Juan Carlos Andresen , Anders Holst

The FTTH (Fiber To The Home) market currently needs new network maintenance technologies that can, economically and effectively, cope with massive fiber plants. However, operating these networks requires adequate means for an effective…

Computational Engineering, Finance, and Science · Computer Science 2016-01-22 Gerson F. M Lima , Edgard Lamounier , Sergio Barcelos , Alexandre Cardoso , Igor Peretta , Willian Muramoto , Flavio Barbara

When testing multiple hypotheses, a suitable error rate should be controlled even in exploratory trials. Conventional methods to control the False Discovery Rate (FDR) assume that all p-values are available at the time point of test…

Methodology · Statistics 2021-12-21 Sonja Zehetmayer , Martin Posch , Franz Koenig

A core strength of knockoff methods is their virtually limitless customizability, allowing an analyst to exploit machine learning algorithms and domain knowledge without threatening the method's robust finite-sample false discovery rate…

Statistics Theory · Mathematics 2021-07-15 Xiao Li , William Fithian

The knockoff-based multiple testing setup of Barber & Candes (2015) for variable selection in multiple regression where sample size is as large as the number of explanatory variables is considered. The method of Benjamini & Hochberg (1995)…

Methodology · Statistics 2021-08-20 Sanat K. Sarkar , Cheng Yong Tang

Time Series Classification (TSC) is essential in fields like medicine, environmental science, and finance, enabling tasks such as disease diagnosis, anomaly detection, and stock price analysis. While machine learning models like Recurrent…

Machine Learning · Computer Science 2024-06-25 Gonzalo Uribarri , Federico Barone , Alessio Ansuini , Erik Fransén

A fundamental problem in the high-dimensional regression is to understand the tradeoff between type I and type II errors or, equivalently, false discovery rate (FDR) and power in variable selection. To address this important problem, we…

Statistics Theory · Mathematics 2020-10-30 Hua Wang , Yachong Yang , Zhiqi Bu , Weijie J. Su

Multiple hypothesis testing has been widely applied to problems dealing with high-dimensional data, e.g., selecting significant variables and controlling the selection error rate. The most prevailing measure of error rate used in the…

Methodology · Statistics 2022-06-07 Xiaoya Sun , Yan Fu

False discovery rates (FDR) are an essential component of statistical inference, representing the propensity for an observed result to be mistaken. FDR estimates should accompany observed results to help the user contextualize the relevance…

Methodology · Statistics 2020-10-12 Megan Hollister Murray , Jeffrey D. Blume

This paper develops a method based on model-X knockoffs to find conditional associations that are consistent across diverse environments, controlling the false discovery rate. The motivation for this problem is that large data sets may…

Methodology · Statistics 2021-06-09 Shuangning Li , Matteo Sesia , Yaniv Romano , Emmanuel Candès , Chiara Sabatti

Selecting important features in high-dimensional survival analysis is critical for identifying confirmatory biomarkers while maintaining rigorous error control. In this paper, we propose a derandomized knockoffs procedure for Cox regression…

Methodology · Statistics 2025-12-15 Rui Liu , Nan Sun