English
Related papers

Related papers: Safe, Always-Valid Alpha-Investing Rules For Doubl…

200 papers

Consider the online testing of a stream of hypotheses where a real--time decision must be made before the next data point arrives. The error rate is required to be controlled at {all} decision points. Conventional \emph{simultaneous testing…

Methodology · Statistics 2020-03-03 Bowen Gang , Wenguang Sun , Weinan Wang

In the online multiple testing problem, p-values corresponding to different null hypotheses are observed one by one, and the decision of whether or not to reject the current hypothesis must be made immediately, after which the next p-value…

Methodology · Statistics 2017-10-03 Aaditya Ramdas , Fanny Yang , Martin J. Wainwright , Michael I. Jordan

In this paper, we propose SAMBA, a novel framework for safe reinforcement learning that combines aspects from probabilistic modelling, information theory, and statistics. Our method builds upon PILCO to enable active exploration using…

A/B tests are typically analyzed via frequentist p-values and confidence intervals; but these inferences are wholly unreliable if users endogenously choose samples sizes by *continuously monitoring* their tests. We define *always valid*…

Statistics Theory · Mathematics 2019-07-18 Ramesh Johari , Leo Pekelis , David J. Walsh

We propose a method that combines the closed testing framework with the concept of safe anytime-valid inference (SAVI) to compute lower confidence bounds for the true discovery proportion in a multiple testing setting. The proposed…

Methodology · Statistics 2026-04-23 Friederike Preusse

Feature selection from a large number of covariates (aka features) in a regression analysis remains a challenge in data science, especially in terms of its potential of scaling to ever-enlarging data and finding a group of scientifically…

Machine Learning · Statistics 2020-02-10 Yiying Fan , Jiayang Sun

Multi-state survival analysis (MSA) uses multi-state models for the analysis of time-to-event data. In medical applications, MSA can provide insights about the complex disease progression in patients. A key challenge in MSA is the accurate…

Machine Learning · Computer Science 2022-07-13 Md Mahmudur Rahman , Sanjay Purushotham

Selecting data for training machine learning models is crucial since large, web-scraped, real datasets contain noisy artifacts that affect the quality and relevance of individual data points. These noisy artifacts will impact model…

Machine Learning · Computer Science 2025-03-20 Samuel Kessler , Tam Le , Vu Nguyen

Financial fraud is the cause of multi-billion dollar losses annually. Traditionally, fraud detection systems rely on rules due to their transparency and interpretability, key features in domains where decisions need to be explained.…

Machine Learning · Computer Science 2024-08-26 João Lucas Martins , João Bravo , Ana Sofia Gomes , Carlos Soares , Pedro Bizarro

Formula alpha mining, which generates predictive signals from financial data, is critical for quantitative investment. Although various algorithmic approaches-such as genetic programming, reinforcement learning, and large language…

Artificial Intelligence · Computer Science 2025-08-20 Hongjun Ding , Binqi Chen , Jinsheng Huang , Taian Guo , Zhengyang Mao , Guoyi Shao , Lutong Zou , Luchen Liu , Ming Zhang

Artificial intelligence enables unprecedented attacks on human cognition, yet cybersecurity remains predominantly device-centric. This paper introduces the "Think First, Verify Always" (TFVA) protocol, which repositions humans as 'Firewall…

Human-Computer Interaction · Computer Science 2025-08-07 Yuksel Aydin

The automated mining of predictive signals, or alphas, is a central challenge in quantitative finance. While Reinforcement Learning (RL) has emerged as a promising paradigm for generating formulaic alphas, existing frameworks are…

Computational Finance · Quantitative Finance 2026-05-20 Binqi Chen , Hongjun Ding , Ning Shen , Jinsheng Huang , Taian Guo , Luchen Liu , Ming Zhang

Always-valid concentration inequalities are increasingly used as performance measures for online statistical learning, notably in the learning of generative models and supervised learning. Such inequality advances the online learning…

Machine Learning · Statistics 2022-11-21 Chi-Hua Wang , Wenjie Li

Online testing procedures assume that hypotheses are observed in sequence, and allow the significance thresholds for upcoming tests to depend on the test statistics observed so far. Some of the most popular online methods include alpha…

Methodology · Statistics 2022-02-11 Aaron Fisher

Many online platforms have deployed anti-fraud systems to detect and prevent fraudulent activities. However, there is usually a gap between the time that a user commits a fraudulent action and the time that the user is suspended by the…

Machine Learning · Computer Science 2018-11-15 Panpan Zheng , Shuhan Yuan , Xintao Wu

Two-sample hypothesis testing for network comparison presents many significant challenges, including: leveraging repeated network observations and known node registration, but without requiring them to operate; relaxing strong structural…

Methodology · Statistics 2024-02-05 Meijia Shao , Dong Xia , Yuan Zhang , Qiong Wu , Shuo Chen

Safe offline reinforcement learning aims to learn policies that maximize cumulative rewards while adhering to safety constraints, using only offline data for training. A key challenge is balancing safety and performance, particularly when…

Machine Learning · Computer Science 2024-12-13 Prajwal Koirala , Zhanhong Jiang , Soumik Sarkar , Cody Fleming

The vulnerability of neural networks to adversarial perturbations has necessitated formal verification techniques that can rigorously certify the quality of neural networks. As the state-of-the-art, branch and bound (BaB) is a…

Machine Learning · Computer Science 2025-07-24 Guanqin Zhang , Kota Fukuda , Zhenya Zhang , H. M. N. Dilum Bandara , Shiping Chen , Jianjun Zhao , Yulei Sui

In the online false discovery rate (FDR) problem, one observes a possibly infinite sequence of $p$-values $P_1,P_2,\dots$, each testing a different null hypothesis, and an algorithm must pick a sequence of rejection thresholds…

Methodology · Statistics 2019-07-12 Aaditya Ramdas , Tijana Zrnic , Martin Wainwright , Michael Jordan

Semi-supervised learning (SSL) enables prediction with limited labels, but high-stakes tabular applications (medical, credit, recidivism) require statistical fairness guarantees. We identify a structural conflict in tabular fair SSL through…

Machine Learning · Computer Science 2026-05-19 Hangchun Liang , Changchun Li
‹ Prev 1 2 3 10 Next ›