中文
相关论文

相关论文: Skilled Mutual Fund Selection: False Discovery Con…

200 篇论文

Given a nonparametric Hidden Markov Model (HMM) with two states, the question of constructing efficient multiple testing procedures is considered, treating one of the states as an unknown null hypothesis. A procedure is introduced, based on…

统计理论 · 数学 2021-01-12 Kweku Abraham , Ismael Castillo , Elisabeth Gassiat

We study unbinned multivariate analysis techniques, based on Statistical Learning, for indirect new physics searches at the LHC in the Effective Field Theory framework. We focus in particular on high-energy $ZW$ production with fully…

高能物理 - 唯象学 · 物理学 2023-01-11 Siyu Chen , Alfredo Glioti , Giuliano Panico , Andrea Wulzer

Federated learning (FL) has received high interest from researchers and practitioners to train machine learning (ML) models for healthcare. Ensuring the trustworthiness of these models is essential. Especially bias, defined as a disparity…

机器学习 · 计算机科学 2023-05-04 Konstantin D. Pandl , Florian Leiser , Scott Thiebes , Ali Sunyaev

Collaborative Filtering (CF) is a widely used technique which allows to leverage past users' preferences data to identify behavioural patterns and exploit them to predict custom recommendations. In this work, we illustrate our review of…

信息检索 · 计算机科学 2022-09-28 Andrea Pinto , Giacomo Camposampiero , Loïc Houmard , Marc Lundwall

Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the…

统计方法学 · 统计学 2022-03-10 Xuebin Zhao , Hong Chen , Yingjie Wang , Weifu Li , Tieliang Gong , Yulong Wang , Feng Zheng

Estimating probability of failure in aerospace systems is a critical requirement for flight certification and qualification. Failure probability estimation involves resolving tails of probability distribution, and Monte Carlo sampling…

数值分析 · 数学 2022-09-22 S. Ashwin Renganathan , Vishwas Rao , Ionel M. Navon

Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery…

机器学习 · 统计学 2025-07-18 Omar Melikechi , David B. Dunson , Jeffrey W. Miller

Conditional independence testing (CIT) is essential for reliable scientific discovery. It prevents spurious findings and enables controlled feature selection. Recent CIT methods have used machine learning (ML) models as surrogates of the…

统计理论 · 数学 2026-02-02 Angel Reyero-Lobo , Bertrand Thirion , Pierre Neuvial

The popularity of penalized regression in high-dimensional data analysis has led to a demand for new inferential tools for these models. False discovery rate control is widely used in high-dimensional hypothesis testing, but has only…

统计方法学 · 统计学 2019-01-24 Ryan Miller , Patrick Breheny

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but developing high-performing models for specialized applications often requires substantial human annotation -- a process that is…

计算与语言 · 计算机科学 2025-07-30 Abhinav Arabelly , Jagrut Nemade , Robert D Nowak , Jifan Zhang

This paper presents a clustering approach that allows for rigorous statistical error control similar to a statistical test. We develop estimators for both the unknown number of clusters and the clusters themselves. The estimators depend on…

统计理论 · 数学 2017-07-13 Michael Vogt , Matthias Schmid

In many modern machine learning applications, the outcome is expensive or time-consuming to collect while the predictor information is easy to obtain. Semi-supervised learning (SSL) aims at utilizing large amounts of `unlabeled' data along…

统计方法学 · 统计学 2017-11-16 Jessica Gronsbell , Tianxi Cai

The performance of a machine learning system is usually evaluated by using i.i.d.\ observations with true labels. However, acquiring ground truth labels is expensive, while obtaining unlabeled samples may be cheaper. Stratified sampling can…

机器学习 · 计算机科学 2019-07-29 Tiancheng Yu , Xiyu Zhai , Suvrit Sra

We investigate a class of methods for selective inference that condition on a selection event. Such methods follow a two-stage process. First, a data-driven (sub)collection of hypotheses is chosen from some large universe of hypotheses.…

统计方法学 · 统计学 2024-04-09 Jelle Goeman , Aldo Solari

Multiple testing has been a popular topic in statistical research. Although vast works have been done, controlling the false discoveries remains a challenging task when the corresponding test statistics are dependent. Various methods have…

统计理论 · 数学 2022-07-05 Meng Mei , Tao Yu , Yuan Jiang

We present a reinforcement-learning (RL) framework for dynamic hedging of equity index option exposures under realistic transaction costs and position limits. We hedge a normalized option-implied equity exposure (one unit of underlying…

投资组合管理 · 定量金融 2025-12-16 Travon Lucius , Christian Koch , Jacob Starling , Julia Zhu , Miguel Urena , Carrie Hu

The most popular multiple testing procedures are stepwise procedures based on $P$-values for individual test statistics. Included among these are the false discovery rate (FDR) controlling procedures of Benjamini--Hochberg [J. Roy. Statist.…

统计理论 · 数学 2009-06-18 Arthur Cohen , Harold B. Sackrowitz , Minya Xu

There has been a growing interest in deep learning-based prognostic and health management (PHM) for building end-to-end maintenance decision support systems, especially due to the rapid development of autonomous systems. However, the low…

机器学习 · 计算机科学 2021-11-02 Taotao Zhou , Enrique Lopez Droguett , Ali Mosleh , Felix T. S. Chan

Contrastive, self-supervised learning (SSL) is used to train a model that predicts cancer type from miRNA, mRNA or RPPA expression data. This model, a pretrained FT-Transformer, is shown to outperform XGBoost and CatBoost, standard…

机器学习 · 计算机科学 2023-11-17 Christian John Hurry , Emma Slade

Current language model training commonly applies multi-task Supervised Fine-Tuning (SFT) using a homogeneous compute budget across all sub-datasets. This approach is fundamentally sub-optimal: heterogeneous learning dynamics cause…

机器学习 · 计算机科学 2026-03-30 Woosung Koh , Jeyoung Jeon , Youngjin Song , Yujin Cheon , Soowon Oh , Jaehyeong Choi , Se-Young Yun