中文
相关论文

相关论文: StarTrek: Combinatorial Variable Selection with Fa…

200 篇论文

We consider a multiple hypothesis testing problem in a sensor network over the joint spatio-temporal domain. The sensor network is modeled as a graph, with each vertex representing a sensor and a signal over time associated with each…

信号处理 · 电气工程与系统科学 2025-01-23 Xingchao Jian , Martin Gölz , Feng Ji , Wee Peng Tay , Abdelhak M. Zoubir

In many large scale multiple testing applications, the hypotheses often have a known graphical structure, such as gene ontology in gene expression data. Exploiting this graphical structure in multiple testing procedures can improve power as…

统计方法学 · 统计学 2018-12-04 Wenge Guo , Gavin Lynch , Joseph P. Romano

With the emergence of graph databases, the task of frequent subgraph discovery has been extensively addressed. Although the proposed approaches in the literature have made this task feasible, the number of discovered frequent subgraphs is…

数据库 · 计算机科学 2013-08-16 Wajdi Dhifli , Mohamed Moussaoui , Rabie Saidi , Engelbert Mephu Nguifo

Hot subdwarf stars are very important for understanding stellar evolution, stellar astrophysics, and binary star systems. Identifying more such stars can help us better understand their statistical distribution, properties, and evolution.…

太阳与恒星天体物理 · 物理学 2022-12-21 Wei Liu , Yude Bu , Xiaoming Kong , Zhenping Yi , Meng Liu

Controlling false discovery rate (FDR) is crucial for variable selection, multiple testing, among other signal detection problems. In literature, there is certainly no shortage of FDR control strategies when selecting individual features,…

统计方法学 · 统计学 2022-04-11 Jingyuan Liu , Ao Sun , Yuan Ke

Feature selection has been an essential step in developing industry-scale deep Click-Through Rate (CTR) prediction systems. The goal of neural feature selection (NFS) is to choose a relatively small subset of features with the best…

机器学习 · 计算机科学 2021-12-08 Lin Guan , Xia Xiao , Ming Chen , Youlong Cheng

Graphical models are popular tools for exploring relationships among a set of variables. The Gaussian graphical model (GGM) is an important class of graphical models, where the conditional dependence among variables is represented by nodes…

统计方法学 · 统计学 2025-05-30 José Á. Sánchez Gómez , Weibin Mo , Junlong Zhao , Yufeng Liu

In modern scientific experiments, we frequently encounter data that have large dimensions, and in some experiments, such high dimensional data arrive sequentially rather than full data being available all at a time. We develop multiple…

统计方法学 · 统计学 2023-06-09 Rahul Roy , Shyamal K. De , Subir Kumar Bhandari

The False Discovery Rate (FDR) method has recently been described by Miller et al (2001), along with several examples of astrophysical applications. FDR is a new statistical procedure due to Benjamini and Hochberg (1995) for controlling the…

天体物理学 · 物理学 2009-11-07 A. M. Hopkins , C. J. Miller , A. J. Connolly , C. Genovese , R. C. Nichol , L. Wasserman

Multivariate statistics are often available as well as necessary in hypothesis tests. We study how to use such statistics to control not only false discovery rate (FDR) but also positive FDR (pFDR) with good power. We show that FDR can be…

统计理论 · 数学 2008-05-21 Zhiyi Chi

There has been recent interest in extending the ideas of False Discovery Rates (FDR) to variable selection in regression settings. Traditionally the FDR in these settings has been defined in terms of the coefficients of the full regression…

统计方法学 · 统计学 2013-02-12 Max Grazier G'Sell , Trevor Hastie , Robert Tibshirani

Community detection is a fundamental task in graph analysis, with methods often relying on fitting models like the Stochastic Block Model (SBM) to observed networks. While many algorithms can accurately estimate SBM parameters when the…

机器学习 · 统计学 2025-06-05 Leonardo Martins Bianco , Christine Keribin , Zacharie Naulet

We consider the task of discovering gene regulatory networks, which are defined as sets of genes and the corresponding transcription factors which regulate their expression levels. This can be viewed as a variable selection problem,…

统计方法学 · 统计学 2014-12-04 Justin Bleich , Adam Kapelner , Edward I. George , Shane T. Jensen

In many fields of science, we observe a response variable together with a large number of potential explanatory variables, and would like to be able to discover which variables are truly associated with the response. At the same time, we…

统计方法学 · 统计学 2015-10-15 Rina Foygel Barber , Emmanuel J. Candès

We introduce DiffKnock, a diffusion-based knockoff framework for high-dimensional feature selection with finite-sample false discovery rate (FDR) control. DiffKnock addresses two key limitations of existing knockoff methods: preserving…

统计方法学 · 统计学 2025-10-03 Heng Ge , Qing Lu

Click Through Rate (CTR) prediction plays an essential role in recommender systems and online advertising. It is crucial to effectively model feature interactions to improve the prediction performance of CTR models. However, existing…

信息检索 · 计算机科学 2023-11-09 Fangye Wang , Hansu Gu , Dongsheng Li , Tun Lu , Peng Zhang , Ning Gu

Thanks to its fine balance between model flexibility and interpretability, the nonparametric additive model has been widely used, and variable selection for this type of model has been frequently studied. However, none of the existing…

统计方法学 · 统计学 2022-01-10 Xiaowu Dai , Xiang Lyu , Lexin Li

Thresholding--the pruning of nodes or edges based on their properties or weights--is an essential preprocessing tool for extracting interpretable structure from complex network data, yet existing methods face several key limitations.…

社会与信息网络 · 计算机科学 2025-10-07 Adam Schroeder , Russell Funk , Jingyi Guan , Taylor Okonek , Lori Ziegelmeier

As datasets grow richer, an important challenge is to leverage the full features in the data to maximize the number of useful discoveries while controlling for false positives. We address this problem in the context of multiple hypotheses…

统计方法学 · 统计学 2017-11-21 Fei Xia , Martin J. Zhang , James Zou , David Tse

Testing composite null hypotheses arises in various applications, such as mediation and replicability analyses. The problem becomes more challenging in high-throughput experiments where tens of thousands of features are examined…

统计方法学 · 统计学 2025-04-29 Pengfei Lyu , Xianyang Zhang , Hongyuan Cao