中文
相关论文

相关论文: Identifying Higher-order Combinations of Binary Fe…

200 篇论文

This paper explores the multiple testing problem for sparse high-dimensional data with binary outcomes. We propose novel empirical Bayes multiple testing procedures based on a spike-and-slab posterior and then evaluate their performance in…

统计理论 · 数学 2025-06-16 Yu-Chien Bo Ning

A fast Bayesian method that seamlessly fuses classification and hypothesis testing via discriminant analysis is developed. Building upon the original discriminant analysis classifier, modelling components are added to identify…

统计方法学 · 统计学 2019-08-28 Weichang Yu , John T. Ormerod , Michael Stewart

Statistical-Relational Model Discovery aims to find statistically relevant patterns in relational data. For example, a relational dependency pattern may stipulate that a user's gender is associated with the gender of their friends. As with…

机器学习 · 计算机科学 2021-10-20 Richard Mar , Oliver Schulte

Large-scale multiple testing under static factor models is widely used to detect sparse signals in high-dimensional data. However, static factor models are arguably too stringent because they ignore serial correlation, which seriously…

统计理论 · 数学 2025-04-04 Xinxin Yang , Lilun Du

We propose a general, modular method for significance testing of groups (or clusters) of variables in a high-dimensional linear model. In presence of high correlations among the covariables, due to serious problems of identifiability, it is…

统计理论 · 数学 2015-02-12 Jacopo Mandozzi , Peter Bühlmann

The balanced incomplete block design (BIBD) problem is a difficult combinatorial problem with a large number of symmetries, which add complexity to its resolution. In this paper, we propose a dual (integer) problem representation that…

神经与进化计算 · 计算机科学 2024-11-05 David Rodríguez Rueda , Carlos Cotta , Antonio J. Fernández-Leiva

In classification problems, especially those that categorize data into a large number of classes, the classes often naturally follow a hierarchical structure. That is, some classes are likely to share similar structures and features. Those…

机器学习 · 计算机科学 2018-07-25 Denali Molitor , Deanna Needell

In many data analyses, each measurement may come with a simple yes/no correction; for example, belonging to one of two populations or being contaminated or not. Ignoring such binary effects may bias the results, while accounting for them…

宇宙学与河外天体物理 · 物理学 2026-05-13 Marcus Högås , Edvard Mörtsell

Stacking, a potent ensemble learning method, leverages a meta-model to harness the strengths of multiple base models, thereby enhancing prediction accuracy. Traditional stacking techniques typically utilize established learning models, such…

机器学习 · 计算机科学 2024-10-31 Wei Wu , Liang Tang , Zhongjie Zhao , Chung-Piaw Teo

A prominent approach to solving combinatorial optimization problems on parallel hardware is Ising machines, i.e., hardware implementations of networks of interacting binary spin variables. Most Ising machines leverage second-order…

Kernel-based hypothesis tests offer a flexible, non-parametric tool to detect high-order interactions in multivariate data, beyond pairwise relationships. Yet the scalability of such tests is limited by the computationally demanding…

统计方法学 · 统计学 2025-06-09 Zhaolu Liu , Robert L. Peach , Mauricio Barahona

Non-probability samples become increasingly popular in survey statistics but may suffer from selection biases that limit the generalizability of results to the target population. We consider integrating a non-probability sample with a…

统计方法学 · 统计学 2019-08-26 Shu Yang , Jae Kwang Kim , Rui Song

This work introduces a novel, simple, and flexible method to quantify irreversibility in generic high-dimensional time series based on the well-known mapping to a binary classification problem. Our approach utilizes gradient boosting for…

统计力学 · 物理学 2025-01-09 Michele Vodret , Cristiano Pacini , Christian Bongiorno

In this paper, we present novel methodologies that incorporate auxiliary variables for multiple hypotheses testing related to the main point of interest while effectively controlling the false discovery rate. When dealing with multiple…

统计方法学 · 统计学 2026-02-23 Seohwa Hwang , Mark Louie Ramos , DoHwan Park , Junyong Park , Johan Lim , Erin Green

In many data exploration tasks it is meaningful to identify groups of attribute interactions that are specific to a variable of interest. For instance, in a dataset where the attributes are medical markers and the variable of interest…

机器学习 · 统计学 2017-03-17 Andreas Henelius , Antti Ukkonen , Kai Puolamäki

In the high dimensional regression analysis when the number of predictors is much larger than the sample size, an important question is to select the important variable which are relevant to the response variable of interest. Variable…

统计方法学 · 统计学 2023-01-09 Pengsheng Ji , Zhigen Zhao

Significant pattern mining, the problem of finding itemsets that are significantly enriched in one class of objects, is statistically challenging, as the large space of candidate patterns leads to an enormous multiple testing problem.…

机器学习 · 统计学 2015-08-25 Felipe Llinares-Lopez , Laetitia Papaxanthos , Dean Bodenham , Karsten Borgwardt

We consider the high-dimensional discriminant analysis problem. For this problem, different methods have been proposed and justified by establishing exact convergence rates for the classification risk, as well as the l2 convergence results…

机器学习 · 统计学 2013-06-28 Mladen Kolar , Han Liu

This paper introduces an innovative method for conducting conditional independence testing in high-dimensional data, facilitating the automated discovery of significant associations within distinct subgroups of a population, all while…

统计方法学 · 统计学 2023-09-19 Matteo Sesia , Tianshu Sun

This work considers the problem of binary classification: given training data $x_1, \dots, x_n$ from a certain population, together with associated labels $y_1,\dots, y_n \in \left\{0,1 \right\}$, determine the best label for an element $x$…

统计理论 · 数学 2016-07-04 Nicolas Garcia Trillos , Ryan Murray