English
Related papers

Related papers: Direct estimation and inference of higher-level co…

200 papers

A novel approach to improve prediction and inference in M-estimation by integrating external information from heterogeneous populations is proposed. Our method leverages joint asymptotics to combine estimates from external and internal…

Methodology · Statistics 2025-09-08 Walter Dempsey , Jeremy M. G. Taylor

Determining the number of factors in high-dimensional factor modeling is essential but challenging, especially when the data are heavy-tailed. In this paper, we introduce a new estimator based on the spectral properties of Spearman sample…

Methodology · Statistics 2024-08-29 Jiaxin Qiu , Zeng Li , Jianfeng Yao

Latent factor models that integrate data from multiple sources/studies or modalities have garnered considerable attention across various disciplines. However, existing methods predominantly focus either on multi-study integration or…

Methodology · Statistics 2025-07-15 Wei Liu , Qingzhi Zhong

Detecting changes in high-dimensional vectors presents significant challenges, especially when the post-change distribution is unknown and time-varying. This paper introduces a novel robust algorithm for correlation change detection in…

Methodology · Statistics 2024-10-07 Assma Alghamdi , Taposh Banerjee , Jayant Rajgopal

Regression models, in which the observed features $X \in \R^p$ and the response $Y \in \R$ depend, jointly, on a lower dimensional, unobserved, latent vector $Z \in \R^K$, with $K< p$, are popular in a large array of applications, and…

Methodology · Statistics 2021-03-04 Xin Bing , Florentina Bunea , Marten Wegkamp

Compositional data arise in many areas of research in the natural and biomedical sciences. One prominent example is in the study of the human gut microbiome, where one can measure the relative abundance of many distinct microorganisms in a…

Methodology · Statistics 2024-04-26 Aaron J. Molstad , Karl Oskar Ekvall , Piotr M. Suder

Estimating individual-level treatment effect from observational data is a fundamental problem in causal inference and has attracted increasing attention in the fields of education, healthcare, and public policy.In this work, we concentrate…

Machine Learning · Computer Science 2025-07-10 Hui Meng , Keping Yang , Xuyu Peng , Bo Zheng

This paper deals with the dimension reduction for high-dimensional time series based on common factors. In particular we allow the dimension of time series $p$ to be as large as, or even larger than, the sample size $n$. The estimation for…

Statistics Theory · Mathematics 2010-06-15 Clifford Lam , Qiwei Yao , Neil Bathia

We consider a binary sequence generated by thresholding a hidden continuous sequence. The hidden variables are assumed to have a compound symmetry covariance structure with a single parameter characterizing the common correlation. We study…

Statistics Theory · Mathematics 2019-09-04 Haolei Weng , Yang Feng

In this paper we derive the optimal linear shrinkage estimator for the high-dimensional mean vector using random matrix theory. The results are obtained under the assumption that both the dimension $p$ and the sample size $n$ tend to…

Statistics Theory · Mathematics 2018-07-17 Taras Bodnar , Ostap Okhrin , Nestor Parolya

It is commonly assumed that a specific testing occasion (task, design, procedure, etc.) provides insights that generalise beyond that occasion. This assumption is infrequently carefully tested in data. We develop a statistically principled…

Applications · Statistics 2020-03-27 Laura Wall , David Gunawan , Scott D. Brown , Minh-Ngoc Tran , Robert Kohn , Guy E. Hawkins

We consider the estimation of approximate factor models for time series data, where strong serial and cross-sectional correlations amongst the idiosyncratic component are present. This setting comes up naturally in many applications, but…

Methodology · Statistics 2019-12-10 Jiahe Lin , George Michailidis

Complex, multivariable systems are often analyzed by grouping their constituent units into components, sometimes referred to as latent features, which afford physical or biological interpretation. However, a priori many different types of…

Disordered Systems and Neural Networks · Physics 2026-05-01 Philipp Fleig , Ilya Nemenman

Factor score estimation in small sample sizes often encounters parameter bias and convergence failures when constructing hierarchical national/sub-national indices. This paper proposes a novel method for hierarchical factor analysis called…

Methodology · Statistics 2025-08-22 Zachary Esses Johnson

Protein aggregation occurs when misfolded or unfolded proteins physically bind together, and can promote the development of various amyloid diseases. This study aimed to construct surrogate models for predicting protein aggregation via…

Quantitative Methods · Quantitative Biology 2023-04-10 Seungpyo Kang , Minseon Kim , Jiwon Sun , Myeonghun Lee , Kyoungmin Min

Long-term causal inference has drawn increasing attention in many scientific domains. Existing methods mainly focus on estimating average long-term causal effects by combining long-term observational data and short-term experimental data.…

Machine Learning · Computer Science 2025-03-04 Weilin Chen , Ruichu Cai , Junjie Wan , Zeqin Yang , José Miguel Hernández-Lobato

Combining data from various sources empowers researchers to explore innovative questions, for example those raised by conducting healthcare monitoring studies. However, the lack of a unique identifier often poses challenges. Record linkage…

Methodology · Statistics 2025-09-16 Kayané Robach , Stéphanie L van der Pas , Mark A van de Wiel , Michel H Hof

In a shotgun proteomics experiment, proteins are the most biologically meaningful output. The success of proteomics studies depends on the ability to accurately and efficiently identify proteins. Many methods have been proposed to…

Quantitative Methods · Quantitative Biology 2012-11-30 Chao Yang , Zengyou He , Weichuan Yu

The current data explosion poses great challenges to the approximate aggregation with an efficiency and accuracy. To address this problem, we propose a novel approach to calculate the aggregation answers with a high accuracy using only a…

Databases · Computer Science 2019-01-23 Shanshan Han , Hongzhi Wang , Jialin Wan , Jianzhong Li

Genomic regions (or loci) displaying outstanding correlation with some environmental variables are likely to be under selection and this is the rationale of recent methods of identifying selected loci and retrieving functional information…

Populations and Evolution · Quantitative Biology 2013-08-13 Gilles Guillot