English
Related papers

Related papers: Large-dimensional Robust Factor Analysis with Grou…

200 papers

This paper studies optimal estimation of large-dimensional nonlinear factor models. The key challenge is that the observed variables are possibly nonlinear functions of some latent variables where the functional forms are left unspecified.…

Statistics Theory · Mathematics 2023-11-14 Yingjie Feng

High-dimensional group inference is an essential part of statistical methods for analysing complex data sets, including hierarchical testing, tests of interaction, detection of heterogeneous treatment effects and inference for local…

Methodology · Statistics 2020-12-01 Zijian Guo , Claude Renaux , Peter Bühlmann , T. Tony Cai

This paper studies the covariance matrix estimation for high-dimensional time series within a new framework that combines low-rank factor and latent variable-specific cluster structures. The popular methods based on assuming the sparse…

Methodology · Statistics 2025-02-25 Dong Li , Xinghao Qiao , Cheng Yu

There are various approaches to graph learning for data clustering, incorporating different spectral and structural constraints through diverse graph structures. Some methods rely on bipartite graph models, where nodes are divided into two…

Machine Learning · Computer Science 2025-05-14 Amirhossein Javaheri , Daniel P. Palomar

Clinical time series data are critical for patient monitoring and predictive modeling. These time series are typically multivariate and often comprise hundreds of heterogeneous features from different data sources. The grouping of features…

Machine Learning · Computer Science 2025-11-12 Fedor Sergeev , Manuel Burger , Polina Leshetkina , Vincent Fortuin , Gunnar Rätsch , Rita Kuznetsova

Latent factor models are widely used to measure unobserved latent traits in social and behavioral sciences, including psychology, education, and marketing. When used in a confirmatory manner, design information is incorporated, yielding…

Methodology · Statistics 2019-06-14 Yunxiao Chen , Xiaoou Li , Siliang Zhang

Algorithms for node clustering typically focus on finding homophilous structure in graphs. That is, they find sets of similar nodes with many edges within, rather than across, the clusters. However, graphs often also exhibit heterophilous…

Machine Learning · Computer Science 2023-08-15 Sudhanshu Chanpuriya , Cameron Musco

Factor analysis has been extensively used to reveal the dependence structures among multivariate variables, offering valuable insight in various fields. However, it cannot incorporate the spatial heterogeneity that is typically present in…

Methodology · Statistics 2024-11-14 Yanxiu Jin , Tomoya Wakayama , Renhe Jiang , Shonosuke Sugasawa

The amount of high-dimensional large-scale RNA sequencing data derived from multiple heterogeneous sources has increased exponentially in biological science. During data collection, significant technical noise or errors may occur. To…

Applications · Statistics 2025-06-24 Xiaolu Jiang , Wei Liu

Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are…

Methodology · Statistics 2018-08-14 Jianqing Fan , Kaizheng Wang , Yiqiao Zhong , Ziwei Zhu

Peer-grouping is used in many sectors for organisational learning, policy implementation, and benchmarking. Clustering provides a statistical, data-driven method for constructing meaningful peer groups, but peer groups must be compatible…

Applications · Statistics 2021-07-14 Daniel William Kennedy , Jessica Cameron , Paul Pao-Yen Wu , Kerrie Mengersen

Crowdsourcing has emerged as an alternative solution for collecting large scale labels. However, the majority of recruited workers are not domain experts, so their contributed labels could be noisy. In this paper, we propose a two-stage…

Methodology · Statistics 2023-09-28 Qi Xu , Yubai Yuan , Junhui Wang , Annie Qu

Considered here are robust subgroup-classifier learning and testing in change-plane regressions with heavy-tailed errors, which can identify subgroups as a basis for making optimal recommendations for individualized treatment. A new…

Methodology · Statistics 2024-08-27 Xu Liu , Jian Huang , Yong Zhou , Xiao Zhang

Statistical learning in high-dimensional spaces is challenging without a strong underlying data structure. Recent advances with foundational models suggest that text and image data contain such hidden structures, which help mitigate the…

Machine Learning · Statistics 2025-02-04 Charles Arnal , Clement Berenfeld , Simon Rosenberg , Vivien Cabannes

Multimodal data, where different types of data are collected from the same subjects, are fast emerging in a large variety of scientific applications. Factor analysis is commonly used in integrative analysis of multimodal data, and is…

Statistics Theory · Mathematics 2021-03-31 Quefeng Li , Lexin Li

In this paper, we set up the theoretical foundations for a high-dimensional functional factor model approach in the analysis of large cross-sections (panels) of functional time series (FTS). We first establish a representation result…

Statistics Theory · Mathematics 2021-04-14 Shahin Tavakoli , Gilles Nisol , Marc Hallin

One potential solution to combat the scarcity of tail observations in extreme value analysis is to integrate information from multiple datasets sharing similar tail properties, for instance, a common extreme value index. In other words, for…

Methodology · Statistics 2025-06-25 Liujun Chen , Marco Oesting , Chen Zhou

Estimation of high-dimensional covariance matrices in latent factor models is an important topic in many fields and especially in finance. Since the number of financial assets grows while the estimation window length remains of limited…

Statistical Finance · Quantitative Finance 2024-07-08 Lucija Žignić , Stjepan Begušić , Zvonko Kostanjčar

We propose a new model selection criterion for mixed effects regression models that is computable when the model is fitted with a two-step method, even when the structure and the distribution of the random effects are unknown. The criterion…

Methodology · Statistics 2018-03-14 Radu V. Craiu , Thierry Duchesne

We propose a combined model, which integrates the latent factor model and the logistic regression model, for the citation network. It is noticed that neither a latent factor model nor a logistic regression model alone is sufficient to…

Machine Learning · Statistics 2019-12-03 Namjoon Suh , Xiaoming Huo , Eric Heim , Lee Seversky