English
Related papers

Related papers: Efficient Estimation Under Data Fusion

200 papers

In this paper we applied data fusion approaches for predicting the final academic performance of university students using multiple-source, multimodal data from blended learning environments. We collected and preprocessed data about…

Computers and Society · Computer Science 2024-03-12 W. Chango , R. Cerezo , C. Romero

Data from both a randomized trial and an observational study are sometimes simultaneously available for evaluating the effect of an intervention. The randomized data typically allows for reliable estimation of average treatment effects but…

Methodology · Statistics 2021-12-01 David Cheng , Tianxi Cai

We present a general framework for using existing data to estimate the efficiency gain from using a covariate-adjusted estimator of a marginal treatment effect in a future randomized trial. We describe conditions under which it is possible…

Methodology · Statistics 2021-05-03 Xiudi Li , Sijia Li , Alex Luedtke

This paper proposes a new approach to multi-sensor data fusion. It suggests that aggregation of data from multiple sensors can be done more efficiently when we consider information about sensors' different characteristics. Similar to most…

Systems and Control · Electrical Eng. & Systems 2019-09-10 Mohammad Amin Ahmad Akhoundi , Ehsan Valavi

Besides the classical motivation of fusing evidence from multiple sources, modern inferential procedures based on randomization, resampling, and data splitting often introduce analyst-generated multiplicity, where aggregating outputs across…

Methodology · Statistics 2026-05-29 Leonardo Cella

Estimating the causal dose-response function is challenging, particularly when data from a single source are insufficient to estimate responses precisely across all exposure levels. To overcome this limitation, we propose a data fusion…

Methodology · Statistics 2025-10-23 Jaewon Lim , Alex Luedtke

We propose a functional accelerated failure time model to characterize effects of both functional and scalar covariates on the time to event of interest, and provide regularity conditions to guarantee model identifiability. For efficient…

Methodology · Statistics 2024-02-09 Changyu Liu , Wen Su , Kin-Yat Liu , Guosheng Yin , Xingqiu Zhao

Data-fusion involves the integration of multiple related datasets. The statistical file-matching problem is a canonical data-fusion problem in multivariate analysis, where the objective is to characterise the joint distribution of a set of…

Methodology · Statistics 2021-04-08 Daniel Ahfock , Saumyadipta Pyne , Geoffrey J. McLachlan

A sequence of social sensors estimate an unknown parameter (modeled as a state of nature) by performing Bayesian Social Learning, and myopically optimize individual reward functions. The decisions of the social sensors contain quantized…

Social and Information Networks · Computer Science 2020-03-31 Sujay Bhatt , Vikram Krishnamurthy

With the development of biomedical science, researchers have increasing access to an abundance of studies focusing on similar research questions. There is a growing interest in the integration of summary information from those studies to…

Methodology · Statistics 2023-11-13 Jianxuan Zang , K. C. G. Chan , Fei Gao

In the age of large and heterogeneous datasets, the integration of information from diverse sources is essential to improve parameter estimation. Multi-task learning offers a powerful approach by enabling simultaneous learning across…

Methodology · Statistics 2025-07-11 Sohom Bhattacharya , Yongzhuo Chen , Muxuan Liang

In semivarying coefficient models for longitudinal/clustered data, usually of primary interest is usually the parametric component which involves unknown constant coefficients. First, we study semiparametric efficiency bound for estimation…

Methodology · Statistics 2015-09-15 Ming-Yen Cheng , Toshio Honda , Jialiang Li

Most prognostic methods require a decent amount of data for model training. In reality, however, the amount of historical data owned by a single organization might be small or not large enough to train a reliable prognostic model. To…

Machine Learning · Statistics 2024-04-11 Madi Arabi , Xiaolei Fang

We address the problem of integrating data from multiple, possibly biased, observational and interventional studies, to eventually compute counterfactuals in structural causal models. We start from the case of a single observational dataset…

Methodology · Statistics 2023-08-01 Marco Zaffalon , Alessandro Antonucci , Rafael Cabañas , David Huber

Nonresponse after probability sampling is a universal challenge in survey sampling, often necessitating adjustments to mitigate sampling and selection bias simultaneously. This study explored the removal of bias and effective utilization of…

Methodology · Statistics 2025-11-13 Kosuke Morikawa , Kenji Beppu , Wataru Aida

We consider the problem of combining data from observational and experimental sources to make causal conclusions. This problem is increasingly relevant, as the modern era has yielded passive collection of massive observational datasets in…

Methodology · Statistics 2020-05-19 Evan Rosenman , Guillaume Basse , Art Owen , Michael Baiocchi

This paper presents a weighted optimization framework that unifies the binary,multi-valued, continuous, as well as mixture of discrete and continuous treatment, under the unconfounded treatment assignment. With a general loss function, the…

Econometrics · Economics 2018-08-20 Chunrong Ai , Oliver Linton , Kaiji Motegi , Zheng Zhang

We provide a unified approach to a method of estimation of the regression parameter in balanced linear models with a structured covariance matrix that combines a high breakdown point and bounded influence with high asymptotic efficiency at…

Statistics Theory · Mathematics 2023-03-22 Hendrik Paul Lopuhaä

Federated learning of causal estimands may greatly improve estimation efficiency by leveraging data from multiple study sites, but robustness to heterogeneity and model misspecifications is vital for ensuring validity. We develop a…

Methodology · Statistics 2023-10-06 Larry Han , Jue Hou , Kelly Cho , Rui Duan , Tianxi Cai

The major sources of abundant data are constantly expanding with the available data collection methodologies in various applications - medical, insurance, scientific, bio-informatics and business. These data sets may be distributed…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-06-24 Aruna Govada , Sanjay K. Sahay