English
Related papers

Related papers: Simultaneous Edit and Imputation for Household Dat…

200 papers

Multivariate time series data for real-world applications typically contain a significant amount of missing values. The dominant approach for classification with such missing values is to impute them heuristically with specific values…

Machine Learning · Computer Science 2023-08-15 SeungHyun Kim , Hyunsu Kim , EungGu Yun , Hwangrae Lee , Jaehun Lee , Juho Lee

Many dynamic microsimulation models have shown their ability to reasonably project detailed population and households using non-data based household formation and dissolution rules. Although, those rules allow modellers to simplify changes…

General Economics · Economics 2020-06-18 Amarin Siripanich , Taha Rashidi

In longitudinal studies, it is not uncommon to make multiple attempts to collect a measurement after baseline. Recording whether these attempts are successful provides useful information for the purposes of assessing missing data…

Methodology · Statistics 2023-05-10 Michael J. Daniels , Minji Lee , Wei Feng

How to include censored data in a statistical analysis is a recur-rent issue in statistics. In multivariate extremes, the dependence structure of large observations can be characterized in terms of a non parametric angular measure, while…

Methodology · Statistics 2014-12-03 Anne Sabourin

Missing values are largely inevitable in gene expression microarray studies. Data sets often have significant omissions due to individuals dropping out of experiments, errors in data collection, image corruptions, and so on. Missing data…

Quantitative Methods · Quantitative Biology 2018-09-18 Marie Li

The US Census Bureau will deliberately corrupt data sets derived from the 2020 US Census, enhancing the privacy of respondents while potentially reducing the precision of economic analysis. To investigate whether this trade-off is…

Econometrics · Economics 2024-02-13 Anish Agarwal , Rahul Singh

In population studies, it is standard to sample data via designs in which the population is divided into strata, with the different strata assigned different probabilities of inclusion. Although there have been some proposals for including…

Methodology · Statistics 2014-09-29 T. Kunihama , A. H. Herring , C. T. Halpern , D. B. Dunson

Mixed data refers to a type of data in which variables can be of multiple types, such as continuous, discrete, or categorical. This data is routinely collected in various fields, including healthcare and social sciences. A common goal in…

Methodology · Statistics 2025-05-22 Mauro Florez , Anna Gottard , Carrie McAdams , Michele Guindani , Marina Vannucci

Imputation of missing attribute values in medical datasets for extracting hidden knowledge from medical datasets is an interesting research topic of interest which is very challenging. One cannot eliminate missing values in medical records.…

Databases · Computer Science 2016-03-11 Yelipe UshaRani , P. Sammulal

Missingness in categorical data is a common problem in various real applications. Traditional approaches either utilize only the complete observations or impute the missing data by some ad hoc methods rather than the true conditional…

Methodology · Statistics 2019-07-12 Chaojie Wang , Linghao Shen , Han Li , Xiaodan Fan

This paper focuses on the problem of hierarchical non-overlapping clustering of a dataset. In such a clustering, each data item is associated with exactly one leaf node and each internal node is associated with all the data items stored in…

Machine Learning · Statistics 2021-05-26 Weipeng Huang , Nishma Laitonjam , Guangyuan Piao , Neil Hurley

Integrating high-dimensional, heterogeneous data from multi-site cohort studies with complex hierarchical structures poses significant feature selection and prediction challenges. We extend the Bayesian Integrative Analysis and Prediction…

Methodology · Statistics 2025-01-30 Aidan Neher , Apostolos Stamenos , Mark Fiecas , Sandra Safo , Thierry Chekouo

Discrete random structures are important tools in Bayesian nonparametrics and the resulting models have proven effective in density estimation, clustering, topic modeling and prediction, among others. In this paper, we consider nested…

Statistics Theory · Mathematics 2018-01-17 Federico Camerlenghi , David B. Dunson , Antonio Lijoi , Igor Prünster , Abel Rodríguez

The advent of Generative Artificial Intelligence (GAI) has heralded an inflection point that changed how society thinks about knowledge acquisition. While GAI cannot be fully trusted for decision-making, it may still provide valuable…

Methodology · Statistics 2025-05-20 Sean O'Hagan , Veronika Ročková

The problem of overdispersed claim counts and mismeasured covariates is common in insurance. On the one hand, the presence of overdispersion in the count data violates the homogeneity assumption, and on the other hand, measurement errors in…

Methodology · Statistics 2023-10-12 Minkun Kim

A Bayesian multivariate model with a structured covariance matrix for multi-way nested data is proposed. This flexible modeling framework allows for positive and for negative associations among clustered observations, and generalizes the…

Methodology · Statistics 2024-08-27 Stef Baas , Richard J. Boucherie , Jean-Paul Fox

Frequentist statistical methods, such as hypothesis testing, are standard practice in papers that provide benchmark comparisons. Unfortunately, these methods have often been misused, e.g., without testing for their statistical test…

Methodology · Statistics 2021-05-18 David Issa Mattos , Jan Bosch , Helena Holmström Olsson

Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…

Statistics Theory · Mathematics 2018-09-03 Irene Dowding , Stefan Haufe

In computational inverse problems, it is common that a detailed and accurate forward model is approximated by a computationally less challenging substitute. The model reduction may be necessary to meet constraints in computing time when…

Methodology · Statistics 2018-02-14 Daniela Calvetti , Matthew M. Dunlop , Erkki Somersalo , Andrew M. Stuart

The vast majority of network datasets contains errors and omissions, although this is rarely incorporated in traditional network analysis. Recently, an increasing effort has been made to fill this methodological gap by developing network…

Social and Information Networks · Computer Science 2018-10-19 Tiago P. Peixoto
‹ Prev 1 4 5 6 7 8 10 Next ›