English
Related papers

Related papers: Considerations for missing data, outliers and tran…

200 papers

High-order parametric models that include terms for feature interactions are applied to various data mining tasks, where ground truth depends on interactions of features. However, with sparse data, the high- dimensional parameters for…

Machine Learning · Computer Science 2018-01-09 Ruocheng Guo , Hamidreza Alvari , Paulo Shakarian

We develop our previous works concerning the identification of the collection of significant factors determining some, in general, non-binary random response variable. Such identification is important, e.g., in biological and medical…

Statistics Theory · Mathematics 2014-06-05 Alexander V. Bulinski , Alexander S. Rakitko

Matching is one of the most widely used causal inference frameworks in observational studies. However, all the existing matching-based causal inference methods are designed for either a single treatment with general treatment types (e.g.,…

Complete randomization allows for consistent estimation of the average treatment effect based on the difference in means of the outcomes without strong modeling assumptions on the outcome-generating process. Appropriate use of the…

Methodology · Statistics 2021-08-03 Anqi Zhao , Peng Ding

It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…

Methodology · Statistics 2023-11-14 Samuel D. Pimentel , Yaxuan Huang

Classical semiparametric inference with missing outcome data is not robust to contamination of the observed data and a single observation can have arbitrarily large influence on estimation of a parameter of interest. This sensitivity is…

Methodology · Statistics 2021-03-02 Eva Cantoni , Xavier de Luna

Modern data science applications often involve complex relational data with dynamic structures. An abrupt change in such dynamic relational data is typically observed in systems that undergo regime changes due to interventions. In such a…

Methodology · Statistics 2024-07-16 Peng Zhao , Anirban Bhattacharya , Debdeep Pati , Bani K. Mallick

We propose VarFA, a variational inference factor analysis framework that extends existing factor analysis models for educational data mining to efficiently output uncertainty estimation in the model's estimated factors. Such uncertainty…

Machine Learning · Statistics 2020-08-18 Zichao Wang , Yi Gu , Andrew Lan , Richard Baraniuk

With many pretreatment covariates and treatment factors, the classical factorial experiment often fails to balance covariates across multiple factorial effects simultaneously. Therefore, it is intuitive to restrict the randomization of the…

Statistics Theory · Mathematics 2018-12-31 Xinran Li , Peng Ding , Donald B. Rubin

Multiple imputation is a straightforward method for handling missing data in a principled fashion. This paper presents an overview of multiple imputation, including important theoretical results and their practical implications for…

Methodology · Statistics 2018-01-15 Jared S. Murray

Advancements in data collection techniques and the heterogeneity of data resources can yield high percentages of missing observations on variables, such as block-wise missing data. Under missing-data scenarios, traditional methods such as…

Methodology · Statistics 2022-05-17 Wei Lan , Xuerong Chen , Tao Zou , Chih-Ling Tsai

Data quality is fundamentally important to ensure the reliability of data for stakeholders to make decisions. In real world applications, such as scientific exploration of extreme environments, it is unrealistic to require raw data…

Artificial Intelligence · Computer Science 2015-10-08 Dongping Fang , Elizabeth Oberlin , Wei Ding , Samuel P. Kounaves

Multivariate regression models and ANOVA are probably the most frequently applied methods of all statistical analyses. We study the case where the predictors are qualitative variables, and the response variable is quantitative. In this…

Applications · Statistics 2021-05-04 Abraham Gutierrez , Sebastian Müller

In some multivariate problems with missing data, pairs of variables exist that are never observed together. For example, some modern biological tools can produce data of this form. As a result of this structure, the covariance matrix is…

Methodology · Statistics 2013-08-13 Max Grazier G'Sell , Shai S. Shen-Orr , Robert Tibshirani

We extend Fisher's randomization test (FRT) to test conditional independence between observed outcomes and treatments given covariates in both randomized experiments and observational studies, with no restriction on the variable type of…

Methodology · Statistics 2025-06-12 Zhen Zhong

This paper studies new tests for the number of latent factors in a large cross-sectional factor model with small time dimension. These tests are based on the eigenvalues of variance-covariance matrices of (possibly weighted) asset returns,…

Econometrics · Economics 2022-10-31 Alain-Philippe Fortin , Patrick Gagliardini , Olivier Scaillet

Factor models are widely used for dimension reduction in the analysis of multivariate data. This is achieved through decomposition of a p x p covariance matrix into the sum of two components. Through a latent factor representation, they can…

Methodology · Statistics 2024-07-01 Sarah Elizabeth Heaps , Ian Hyla Jermyn

Outliers widely occur in big-data applications and may severely affect statistical estimation and inference. In this paper, a framework of outlier-resistant estimation is introduced to robustify an arbitrarily given loss function. It has a…

Methodology · Statistics 2023-04-20 Yiyuan She , Zhifeng Wang , Jiahui Shen

Quantile Factor Models (QFM) represent a new class of factor models for high-dimensional panel data. Unlike Approximate Factor Models (AFM), where only location-shifting factors can be extracted, QFM also allow to recover unobserved factors…

Econometrics · Economics 2020-09-24 Liang Chen , Juan Jose Dolado , Jesus Gonzalo

We develop factor copula models for analysing the dependence among mixed continuous and discrete responses. Factor copula models are canonical vine copulas that involve both observed and latent variables, hence they allow tail, asymmetric…

Methodology · Statistics 2020-11-18 Sayed H. Kadhem , Aristidis K. Nikoloulopoulos
‹ Prev 1 8 9 10 Next ›