中文
相关论文

相关论文: Considerations for missing data, outliers and tran…

200 篇论文

Traditional analysis of variance (ANOVA) software allows researchers to test for the significance of main effects in the presence of interactions without exposure to the details of how the software encodes main effects and interactions to…

统计方法学 · 统计学 2018-01-16 Roger Levy

Randomization inference is a widely-used and appealing approach for analyzing treatment effects in randomized experiments, as it is finite-sample valid and does not require any distributional assumptions. However, naive application of…

计量经济学 · 经济学 2026-05-12 Xinran Li , Peizan Sheng , Zeyang Yu

We consider the design and analysis of multi-factor experiments using fractional factorial and incomplete designs within the potential outcome framework. These designs are particularly useful when limited resources make running a full…

统计方法学 · 统计学 2022-01-31 Nicole E. Pashley , Marie-Abele C. Bind

We propose a new method to impute missing values in mixed datasets. It is based on a principal components method, the factorial analysis for mixed data, which balances the influence of all the variables that are continuous and categorical…

应用统计 · 统计学 2013-02-20 Vincent Audigier , François Husson , Julie Josse

Economists are blessed with a wealth of data for analysis, but more often than not, values in some entries of the data matrix are missing. Various methods have been proposed to handle missing observations in a few variables. We exploit the…

计量经济学 · 经济学 2022-02-02 Ercument Cahan , Jushan Bai , Serena Ng

Population means and standard deviations are the most common estimands to quantify effects in factorial layouts. In fact, most statistical procedures in such designs are built towards inferring means or contrasts thereof. For more robust…

统计理论 · 数学 2020-03-30 Marc Ditzhaus , Roland Fried , Markus Pauly

Factor modeling is an essential tool for exploring intrinsic dependence structures among high-dimensional random variables. Much progress has been made for estimating the covariance matrix from a high-dimensional factor model. However, the…

统计理论 · 数学 2016-10-26 Quefeng Li , Guang Cheng , Jianqing Fan , Yuyan Wang

Many statistical analyses involve the comparison of multiple data sets collected under different conditions in order to identify the difference in the underlying distributions. A common challenge in multi-sample comparison is the presence…

统计方法学 · 统计学 2016-04-07 Li Ma , Jacopo Soriano

Data fusion techniques integrate information from heterogeneous data sources to improve learning, generalization, and decision making across data sciences. In causal inference, these methods leverage rich observational data to improve…

统计方法学 · 统计学 2025-06-02 Quinn Lanners , Cynthia Rudin , Alexander Volfovsky , Harsh Parikh

Missing data can lead to inefficiencies and biases in analyses, in particular when data are missing not at random (MNAR). It is thus vital to understand and correctly identify the missing data mechanism. Recovering missing values through a…

统计方法学 · 统计学 2022-12-08 Jack Noonan , Adetola Adedamola Adediran , Robin Mitra , Stefanie Biedermann

Non-negative matrix factorization (NMF) has previously been shown to be a useful decomposition for multivariate data. We interpret the factorization in a new way and use it to generate missing attributes from test data. We provide a joint…

数值分析 · 计算机科学 2010-07-05 Mithun Das Gupta

Consider a regression or some regression-type model for a certain response variable where the linear predictor includes an ordered factor among the explanatory variables. The inclusion of a factor of this type can take place is a few…

统计方法学 · 统计学 2023-11-27 Adelchi Azzalini

The goal of this paper is to design a causal inference method accounting for complex interactions between causal factors. The proposed method relies on a category theoretical reformulation of the definitions of dependent variables,…

统计理论 · 数学 2020-06-16 Rémy Tuyéras

Factor analysis is a critical component of high dimensional biological data analysis. However, modern biological data contain two key features that irrevocably corrupt existing methods. First, these data, which include longitudinal,…

统计方法学 · 统计学 2020-09-24 Chris McKennan

Multivariate data are typically represented by a rectangular matrix (table) in which the rows are the objects (cases) and the columns are the variables (measurements). When there are many variables one often reduces the dimension by…

统计方法学 · 统计学 2021-01-13 Mia Hubert , Peter J. Rousseeuw , Wannes Van den Bossche

We study a general factor analysis framework where the $n$-by-$p$ data matrix is assumed to follow a general exponential family distribution entry-wise. While this model framework has been proposed before, we here further relax its…

统计方法学 · 统计学 2025-12-02 Liang Wang , Luis Carvalho

Design-based causal inference, also known as randomization-based or finite-population causal inference, is one of the most widely used causal inference frameworks, largely due to the merit that its validity can be guaranteed by study design…

统计方法学 · 统计学 2025-05-27 Siyu Heng , Jiawei Zhang , Yang Feng

The paper considers linear regression problems where the number of predictor variables is possibly larger than the sample size. The basic motivation of the study is to combine the points of view of model selection and functional regression…

统计理论 · 数学 2012-02-24 Alois Kneip , Pascal Sarda

This paper studies inference in two-stage randomized experiments under covariate-adaptive randomization. In the initial stage of this experimental design, clusters (e.g., households, schools, or graph partitions) are stratified and randomly…

计量经济学 · 经济学 2026-01-16 Jizhou Liu

Factor analysis (FA) and principal component analysis (PCA) are popular statistical methods for summarizing and explaining the variability in multivariate datasets. By default, FA and PCA assume the number of components or factors to be…

统计方法学 · 统计学 2022-05-17 Chetkar Jha , Ian Barnett