中文
相关论文

相关论文: Considerations for missing data, outliers and tran…

200 篇论文

Missing data is a common challenge when analyzing epidemiological data, and imputation is often used to address this issue. Here, we investigate the scenario where a covariate used in an analysis has missingness and will be imputed. There…

统计方法学 · 统计学 2024-03-04 Lucy D'Agostino McGowan , Sarah C. Lotspeich , Staci A. Hepler

In regression analysis of multivariate data, it is tacitly assumed that response and predictor variables in each observed response-predictor pair correspond to the same entity or unit. In this paper, we consider the situation of "permuted…

统计理论 · 数学 2017-11-17 Martin Slawski , Emanuel Ben-David

Observational data is increasingly used as a means for making individual-level causal predictions and intervention recommendations. The foremost challenge of causal inference from observational data is hidden confounding, whose presence…

机器学习 · 统计学 2018-10-30 Nathan Kallus , Aahlad Manas Puli , Uri Shalit

Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent…

机器学习 · 统计学 2024-02-26 Kartik Ahuja , Divyat Mahajan , Yixin Wang , Yoshua Bengio

We give an expository review of applications of computational algebraic statistics to design and analysis of fractional factorial experiments based on our recent works. For the purpose of design, the techniques of Gr\"obner bases and…

统计方法学 · 统计学 2012-04-09 Satoshi Aoki , Akimichi Takemura

In many domains such as healthcare or finance, data often come in different assays or measurement modalities, with features in each assay having a common theme. Simply concatenating these assays together and performing prediction can be…

统计方法学 · 统计学 2018-07-17 J. Kenneth Tay , Robert Tibshirani

Factor analysis (FA) is a statistical tool for studying how observed variables with some mutual dependences can be expressed as functions of mutually independent unobserved factors, and it is widely applied throughout the psychological,…

机器学习 · 统计学 2023-06-01 Alex Markham , Mingyu Liu , Bryon Aragam , Liam Solus

In this paper, a scale mixture of Normal distributions model is developed for classification and clustering of data having outliers and missing values. The classification method, based on a mixture model, focuses on the introduction of…

机器学习 · 统计学 2017-11-23 G. Revillon , A. Djafari , C. Enderli

This paper addresses one of the most prevalent problems encountered by political scientists working with difference-in-differences (DID) design: missingness in panel data. A common practice for handling missing data, known as complete case…

统计方法学 · 统计学 2024-12-02 Sooahn Shin

Item nonresponse is frequently encountered in practice. Ignoring missing data can lose efficiency and lead to misleading inference. Fractional imputation is a frequentist approach of imputation for handling missing data. However, the…

统计方法学 · 统计学 2018-09-18 Hejian Sang , Jae Kwang Kim

Integrative analysis of datasets generated by multiple cohorts is a widely-used approach for increasing sample size, precision of population estimators, and generalizability of analysis results in epidemiological studies. However, often…

Latent or unobserved phenomena pose a significant difficulty in data analysis as they induce complicated and confounding dependencies among a collection of observed variables. Factor analysis is a prominent multivariate statistical modeling…

统计方法学 · 统计学 2020-06-22 Armeen Taeb , Venkat Chandrasekaran

Confirmatory Factor Analysis (CFA) is a particular form of factor analysis, most commonly used in social research. In confirmatory factor analysis, the researcher first develops a hypothesis about what factors they believe are underlying…

应用统计 · 统计学 2019-05-15 Rui Portocarrero Sarmento , Vera Costa

Causal inference starts with a simple idea: compare groups that differ by treatment, not much else. Traditionally, similar groups are constructed using only observed covariates; however, it remains a long-standing challenge to incorporate…

统计方法学 · 统计学 2025-11-21 Ying Jin , José Zubizarreta

Causal inference quantifies cause-effect relationships by estimating counterfactual parameters from data. This entails using \emph{identification theory} to establish a link between counterfactual parameters of interest and distributions…

机器学习 · 统计学 2020-04-17 Jaron J. R. Lee , Ilya Shpitser

Data analyses typically rely upon assumptions about missingness mechanisms that lead to observed versus missing data. When the data are missing not at random, direct assumptions about the missingness mechanism, and indirect assumptions…

统计方法学 · 统计学 2016-03-22 Alexander M Franks , Edoardo M Airoldi , Donald B Rubin

Factor analysis is over a century old, but it is still problematic to choose the number of factors for a given data set. The scree test is popular but subjective. The best performing objective methods are recommended on the basis of…

统计方法学 · 统计学 2015-11-12 A. B. Owen , J. Wang

Estimations and evaluations of the main patterns of time series data in groups benefit large amounts of applications in various fields. Different from the classical auto-correlation time series analysis and the modern neural networks…

应用统计 · 统计学 2022-03-29 Rongjiao Ji , Alessandra Micheletti , Nataša Krklec Jerinkić , Zoranka Desnica

This paper considers the estimation and inference of the low-rank components in high-dimensional matrix-variate factor models, where each dimension of the matrix-variates ($p \times q$) is comparable to or greater than the number of…

统计理论 · 数学 2022-10-20 Elynn Y. Chen , Jianqing Fan

Missing data are ubiquitous in the era of big data and, if inadequately handled, are known to lead to biased findings and have deleterious impact on data-driven decision makings. To mitigate its impact, many missing value imputation methods…

机器学习 · 计算机科学 2021-10-26 Yiliang Zhang , Qi Long