中文
相关论文

相关论文: Multiple hypothesis testing adjusted for latent va…

200 篇论文

The exponential growth of scientific knowledge has made the automated generation of scientific hypotheses that combine novelty, feasibility, and research value a core challenge. Existing methods based on large language models fail to…

人工智能 · 计算机科学 2025-08-05 Shiyang Duan , Yuan Tian , Qi Bing , Xiaowei Shao

In this paper we consider the problem of multiple testing when the hypotheses are dependent. In most of the existing literature, either Bayesian or non-Bayesian, the decision rules mainly focus on the validity of the test procedure rather…

统计方法学 · 统计学 2018-07-10 Noirrit K. Chandra , Sourabh Bhattacharya

A two-sample hypothesis test is a statistical procedure used to determine whether the distributions generating two samples are identical. We consider the two-sample testing problem in a new scenario where the sample measurements (or sample…

In a high dimensional regression setting in which the number of variables ($p$) is much larger than the sample size ($n$), the number of possible two-way interactions between the variables is immense. If the number of variables is in the…

统计方法学 · 统计学 2024-06-26 Marianne A Jonker , Luc van Schijndel , Eric Cator

We consider the high-dimensional sparse linear regression problem of accurately estimating a sparse vector using a small number of linear measurements that are contaminated by noise. It is well known that the standard cadre of…

统计理论 · 数学 2014-02-25 Divyanshu Vats , Richard G. Baraniuk

Multiple testing problems arising in modern scientific applications can involve simultaneously testing thousands or even millions of hypotheses, with relatively few true signals. In this paper, we consider the multiple testing problem where…

统计方法学 · 统计学 2016-06-28 Ang Li , Rina Foygel Barber

In this paper, we present novel methodologies that incorporate auxiliary variables for multiple hypotheses testing related to the main point of interest while effectively controlling the false discovery rate. When dealing with multiple…

统计方法学 · 统计学 2026-02-23 Seohwa Hwang , Mark Louie Ramos , DoHwan Park , Junyong Park , Johan Lim , Erin Green

Finding statistically significant interactions between binary variables is computationally and statistically challenging in high-dimensional settings, due to the combinatorial explosion in the number of hypotheses. Terada et al. recently…

机器学习 · 统计学 2014-07-07 Felipe Llinares , Mahito Sugiyama , Karsten M. Borgwardt

Hierarchical learning models, such as mixture models and Bayesian networks, are widely employed for unsupervised learning tasks, such as clustering analysis. They consist of observable and hidden variables, which represent the given data…

机器学习 · 统计学 2018-01-08 Keisuke Yamazaki

High-dimensional tests are applied to find relevant sets of variables and relevant models. If variables are selected by analyzing the sums of products matrices and a corresponding mean-value test is performed, there is the danger that the…

统计方法学 · 统计学 2012-02-10 Juergen Laeuter , Maciej Rosolowski , Ekkehard Glimm

It is common to show the confidence intervals or $p$-values of selected features, or predictor variables in regression, but they often involve selection bias. The selective inference approach solves this bias by conditioning on the…

统计方法学 · 统计学 2022-06-02 Yoshikazu Terada , Hidetoshi Shimodaira

A large number of recent genome-wide association studies (GWASs) for complex phenotypes confirm the early conjecture for polygenicity, suggesting the presence of large number of variants with only tiny or moderate effects. However, due to…

基因组学 · 定量生物学 2018-05-01 Mingwei Dai , Xiang Wan , Hao Peng , Yao Wang , Yue Liu , Jin Liu , Zongben Xu , Can Yang

Suppose that at any stage of a statistical experiment a control variable $X$ that affects the distribution of the observed data $Y$ at this stage can be used. The distribution of $Y$ depends on some unknown parameter $\theta$, and we…

统计理论 · 数学 2009-12-23 Andrey Novikov

Language models have been shown to perform remarkably well on a wide range of natural language processing tasks. In this paper, we propose LEAP, a novel system that uses language models to perform multi-step logical reasoning and…

计算与语言 · 计算机科学 2023-11-08 Hongyu Zhao , Kangrui Wang , Mo Yu , Hongyuan Mei

We present generalized additive latent and mixed models (GALAMMs) for analysis of clustered data with responses and latent variables depending smoothly on observed variables. A scalable maximum likelihood estimation algorithm is proposed,…

统计方法学 · 统计学 2023-03-29 Øystein Sørensen , Anders M. Fjell , Kristine B. Walhovd

Large Reasoning Models (LRMs) have the ability to self-correct even when they make mistakes in their reasoning paths. However, our study reveals that when the reasoning process starts with a short but poor beginning, it becomes difficult…

计算与语言 · 计算机科学 2025-05-13 Tongxu Luo , Wenyu Du , Jiaxi Bi , Stephen Chung , Zhengyang Tang , Hao Yang , Min Zhang , Benyou Wang

Parameter estimates for associated genetic variants, report ed in the initial discovery samples, are often grossly inflated compared to the values observed in the follow-up replication samples. This type of bias is a consequence of the…

应用统计 · 统计学 2011-04-15 Lizhen Xu , Radu V. Craiu , Lei Sun

Motivated by genetic association studies of pleiotropy, we propose here a Bayesian latent variable approach to jointly study multiple outcomes or phenotypes. The proposed method models both continuous and binary phenotypes, and it accounts…

应用统计 · 统计学 2012-11-08 Lizhen Xu , Radu V. Craiu , Lei Sun

Gene expression and phenotype association can be affected by potential unmeasured confounders from multiple sources, leading to biased estimates of the associations. Since genetic variants largely explain gene expression variations, they…

统计方法学 · 统计学 2019-10-23 Jiarui Lu , Hongzhe Li

Addressing selection bias in latent variable causal discovery is important yet underexplored, largely due to a lack of suitable statistical tools: While various tools beyond basic conditional independencies have been developed to handle…

机器学习 · 计算机科学 2025-12-15 Haoyue Dai , Yiwen Qiu , Ignavier Ng , Xinshuai Dong , Peter Spirtes , Kun Zhang