English
Related papers

Related papers: Data-Adaptive Integration With Summary Data

200 papers

In clinical settings, we often face the challenge of building prediction models based on small observational data sets. For example, such a data set might be from a medical center in a multi-center study. Differences between centers might…

In time series analysis, statistics based on collections of estimators computed from sub-samples play a crucial role in an increasing variety of important applications. Proving results about the joint asymptotic distribution of such…

Statistics Theory · Mathematics 2013-05-27 Stanislav Volgushev , Xiaofeng Shao

A common approach to synthetic data is to sample from a fitted model. We show that under general assumptions, this approach results in a sample with inefficient estimators and whose joint distribution is inconsistent with the true…

Statistics Theory · Mathematics 2026-02-18 Jordan Awan , Zhanrui Cai

Incorporating the auxiliary information into the survey estimation is a fundamental problem in survey sampling. Calibration weighting is a popular tool for incorporating the auxiliary information. The calibration weighting method of Deville…

Methodology · Statistics 2025-05-13 Yonghyun Kwon , Jae Kwang Kim , Yumou Qiu

Obtaining causally interpretable meta-analysis results is challenging when there are differences in the distribution of effect modifiers between eligible trials. To overcome this, recent work on transportability methods has considered…

Methodology · Statistics 2025-03-10 Tat-Thang Vo , Tran Trong Khoi Le , Sivem Afach , Stijn Vansteelandt

Low-order thermoacoustic models are qualitatively correct, but they are typically quantitatively inaccurate. We propose a time-domain bias-aware method to make qualitatively low--order models quantitatively (more) accurate. First, we…

Fluid Dynamics · Physics 2022-11-10 Andrea Nóvoa , Luca Magri

Recent years have experienced increasing utilization of complex machine learning models across multiple sources of data to inform more generalizable decision-making. However, distribution shifts across data sources and privacy concerns…

Methodology · Statistics 2024-05-16 Yi Liu , Alexander W. Levis , Sharon-Lise Normand , Larry Han

The entropic risk measure is widely used in high-stakes decision-making across economics, management science, finance, and safety-critical control systems because it captures tail risks associated with uncertain losses. However, when data…

Optimization and Control · Mathematics 2026-01-05 Utsav Sadana , Erick Delage , Angelos Georghiou

Randomized controlled trials (RCTs) are widely regarded as the gold standard for causal inference in biomedical research. For instance, when estimating the average treatment effect on the treated (ATT), a doubly robust estimation procedure…

Methodology · Statistics 2025-09-26 Chi-Shian Dai , Chao Ying , Yang Ning , Jiwei Zhao

Nonresponse after probability sampling is a universal challenge in survey sampling, often necessitating adjustments to mitigate sampling and selection bias simultaneously. This study explored the removal of bias and effective utilization of…

Methodology · Statistics 2025-11-13 Kosuke Morikawa , Kenji Beppu , Wataru Aida

It is becoming increasingly common for researchers to consider incorporating external information from large studies to improve the accuracy of statistical inference instead of relying on a modestly sized dataset collected internally. With…

Methodology · Statistics 2021-07-20 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

A challenge arising from the local Bayesian assimilation of data in an atmospheric flow simulation is the imbalances it may introduce. Acoustic fast-mode imbalances of the order of the slower dynamics can be negated by employing a blended…

Numerical Analysis · Mathematics 2024-03-11 Ray Chew , Tommaso Benacchio , Gottfried Hastermann , Rupert Klein

We aim to make inferences about a smooth, finite-dimensional parameter by fusing data from multiple sources together. Previous works have studied the estimation of a variety of parameters in similar data fusion settings, including in the…

Methodology · Statistics 2025-02-03 Sijia Li , Alex Luedtke

RNA-sequencing has revolutionized biomedical research and, in particular, our ability to study gene alternative splicing. The problem has important implications for human health, as alternative splicing may be involved in malfunctions at…

Applications · Statistics 2015-12-11 David Rossell , Camille Stephan-Otto Attolini , Manuel Kroiss , Almond Stöcker

Two important considerations in clinical research studies are proper evaluations of internal and external validity. While randomized clinical trials can overcome several threats to internal validity, they may be prone to poor external…

Methodology · Statistics 2022-07-19 Kevin P. Josey , Fan Yang , Debashis Ghosh , Sridharan Raghavan

Repeated use of a data sample via adaptively chosen queries can rapidly lead to overfitting, wherein the empirical evaluation of queries on the sample significantly deviates from their mean with respect to the underlying data distribution.…

Machine Learning · Computer Science 2024-04-26 Moshe Shenfeld , Katrina Ligett

Data aggregation, also known as meta analysis, is widely used to combine knowledge on parameters shared in common (e.g., average treatment effect) between multiple studies. In this paper, we introduce an attractive data aggregation scheme…

Methodology · Statistics 2023-05-10 Snigdha Panigrahi , Jingshen Wang , Xuming He

Subsampling from a large data set is useful in many supervised learning contexts to provide a global view of the data based on only a fraction of the observations. Diverse (or space-filling) subsampling is an appealing subsampling approach…

Methodology · Statistics 2023-11-27 Boyang Shang , Daniel W. Apley , Sanjay Mehrotra

Domain shifts are ubiquitous in machine learning, and can substantially degrade a model's performance when deployed to real-world data. To address this, distribution alignment methods aim to learn feature representations which are invariant…

Machine Learning · Computer Science 2024-10-08 Andrea Napoli , Paul White

An aggregate data meta-analysis is a statistical method that pools the summary statistics of several selected studies to estimate the outcome of interest. When considering a continuous outcome, typically each study must report the same…

Methodology · Statistics 2022-06-22 Sean McGrath , XiaoFei Zhao , Zhi Zhen Qin , Russell Steele , Andrea Benedetti