中文
相关论文

相关论文: Data Integration by combining big data and survey …

200 篇论文

The increasing application of social and human-enabled systems in people's daily life from one side and from the other side the fast growth of mobile and smart phones technologies have resulted in generating tremendous amount of data, also…

人机交互 · 计算机科学 2016-04-19 Mohammad Allahbakhsh , Saeed Arbabi , Hamid-Reza Motahari-Nezhad , Boualem Benatallah

Demographic models built from genetic data play important roles in illuminating prehistorical events and serving as null models in genome scans for selection. We introduce an inference method based on the joint frequency spectrum of genetic…

种群与进化 · 定量生物学 2010-05-10 Ryan N. Gutenkunst , Ryan D. Hernandez , Scott H. Williamson , Carlos D. Bustamante

Surveys are commonly used to facilitate research in epidemiology, health, and the social and behavioral sciences. Often, these surveys are not simple random samples, and respondents are given weights reflecting their probability of…

统计方法学 · 统计学 2024-08-20 Adway S. Wadekar , Jerome P. Reiter

An increasing number of large-scale multi-modal research initiatives has been conducted in the typically developing population, as well as in psychiatric cohorts. Missing data is a common problem in such datasets due to the difficulty of…

Data-driven risk analysis involves the inference of probability distributions from measured or simulated data. In the case of a highly reliable system, such as the electricity grid, the amount of relevant data is often exceedingly limited,…

统计方法学 · 统计学 2017-07-11 Simon H. Tindemans , Goran Strbac

With the ubiquitous availability of unstructured data, growing attention is paid as how to adjust for selection bias in such non-probability samples. The majority of the robust estimators proposed by prior literature are either fully or…

统计方法学 · 统计学 2022-04-08 Ali Rafei , Michael R. Elliott , Carol A. C. Flannagan

Missing data is an universal problem in statistics. We develop a unified framework for estimating parameters defined by general estimating equations under a missing-at-random (MAR) mechanism, based on generalized entropy calibration…

统计方法学 · 统计学 2026-03-31 Mst Moushumi Pervin , Hengfang Wang , Jae Kwang Kim

Class imbalance and distributional differences in large datasets present significant challenges for classification tasks machine learning, often leading to biased models and poor predictive performance for minority classes. This work…

机器学习 · 统计学 2024-12-20 Alex Mak , Shubham Sahoo , Shivani Pandey , Yidan Yue , Linglong Kong

In population studies, it is standard to sample data via designs in which the population is divided into strata, with the different strata assigned different probabilities of inclusion. Although there have been some proposals for including…

统计方法学 · 统计学 2014-09-29 T. Kunihama , A. H. Herring , C. T. Halpern , D. B. Dunson

Increasing nonresponse rates and the cost of data collection are two pressing problems encountered in traditional probability surveys. The proliferation of inexpensive data from web surveys stimulates interest in statistical techniques for…

统计方法学 · 统计学 2019-12-31 Vladislav Beresovsky

The National Health and Nutrition Examination Survey (NHANES) studies the nutritional and health status over the whole U.S. population with comprehensive physical examinations and questionnaires. However, survey data analyses become…

统计方法学 · 统计学 2019-08-06 Xiaojun Mao , Zhonglei Wang , Shu Yang

If part of a population is hidden but two or more sources are available that each cover parts of this population, dual- or multiple-system(s) estimation can be applied to estimate this population. For this it is common to use the log-linear…

统计方法学 · 统计学 2023-11-06 Daan B. Zult , Peter G. M. van der Heijden , Bart F. M. Bakker

Multiple imputation (MI) has become popular for analyses with missing data in medical research. The standard implementation of MI is based on the assumption of data being missing at random (MAR). However, for missing data generated by…

统计方法学 · 统计学 2019-01-03 Tra My Pham , James R Carpenter , Tim P Morris , Angela M Wood , Irene Petersen

Item nonresponse is frequently encountered in practice. Ignoring missing data can lose efficiency and lead to misleading inference. Fractional imputation is a frequentist approach of imputation for handling missing data. However, the…

统计方法学 · 统计学 2018-09-18 Hejian Sang , Jae Kwang Kim

When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our…

统计方法学 · 统计学 2024-04-30 Shirley Mathur , Yajuan Si , Jerome P. Reiter

We propose a distributed method for simultaneous inference for datasets with sample size much larger than the number of covariates, i.e., N >> p, in the generalized linear models framework. When such datasets are too big to be analyzed…

统计方法学 · 统计学 2020-07-23 Lu Tang , Ling Zhou , Peter X. -K. Song

The importance of exploring a potential integration among surveys has been acknowledged in order to enhance effectiveness and minimize expenses. In this work, we employ the alignment method to combine information from two different surveys…

统计方法学 · 统计学 2024-04-09 Vasilis Chasiotis , Dimitris Karlis

Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…

统计理论 · 数学 2018-09-03 Irene Dowding , Stefan Haufe

Multiple imputation is a highly recommended technique to deal with missing data, but the application to longitudinal datasets can be done in multiple ways. When a new wave of longitudinal data arrives, we can treat the combined data of…

统计方法学 · 统计学 2026-05-18 X. M. Kavelaars , S. van Buuren , J. R. van Ginkel

A common problem faced by statistical institutes is that data may be missing from collected data sets. The typical way to overcome this problem is to impute the missing data. The problem of imputing missing data is complicated by the fact…

应用统计 · 统计学 2014-01-09 Jeroen Pannekoek , Natalie Shlomo , Ton De Waal