中文
相关论文

相关论文: Towards a Unified Theory for Semiparametric Data F…

200 篇论文

Many statistical estimands of interest (e.g., in regression or causality) are functions of the joint distribution of multiple random variables. But in some applications, data is not available that measures all random variables on each…

统计方法学 · 统计学 2025-02-11 Yicong Jiang , Lucas Janson

We aim to make inferences about a smooth, finite-dimensional parameter by fusing data from multiple sources together. Previous works have studied the estimation of a variety of parameters in similar data fusion settings, including in the…

统计方法学 · 统计学 2025-02-03 Sijia Li , Alex Luedtke

Suppose one is interested in estimating causal effects in the presence of potentially unmeasured confounding with the aid of a valid instrumental variable. This paper investigates the problem of making inferences about the average treatment…

统计方法学 · 统计学 2020-12-15 BaoLuo Sun , Wang Miao

Causal inference across multiple data sources offers a promising avenue to enhance the generalizability and replicability of scientific findings. However, data integration methods for time-to-event outcomes, common in biomedical research,…

统计方法学 · 统计学 2025-05-16 Yi Liu , Alexander W. Levis , Ke Zhu , Shu Yang , Peter B. Gilbert , Larry Han

We introduce a new data fusion method that utilizes multiple data sources to estimate a smooth, finite-dimensional parameter. Most existing methods only make use of fully aligned data sources that share common conditional distributions of…

统计方法学 · 统计学 2025-04-30 Sijia Li , Peter B. Gilbert , Rui Duan , Alex Luedtke

Suppose we have individual data from an internal study and various summary statistics from relevant external studies. External summary statistics have the potential to improve statistical inference for the internal population; however, it…

统计方法学 · 统计学 2026-02-06 Wenjie Hu , Ruoyu Wang , Wei Li , Wang Miao

Data-fusion involves the integration of multiple related datasets. The statistical file-matching problem is a canonical data-fusion problem in multivariate analysis, where the objective is to characterise the joint distribution of a set of…

统计方法学 · 统计学 2021-04-08 Daniel Ahfock , Saumyadipta Pyne , Geoffrey J. McLachlan

Fusion learning refers to synthesizing inferences from multiple sources or studies to provide more effective inference and prediction than from any individual source or study alone. Most existing methods for synthesizing inferences rely on…

统计方法学 · 统计学 2020-11-16 Dungang Liu , Regina Y. Liu , Minge Xie

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

统计方法学 · 统计学 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

This paper proposes a new theory and methodology to tackle the problem of unifying distributed analyses and inferences on shared parameters from multiple sources, into a single coherent inference. This surprisingly challenging problem…

统计方法学 · 统计学 2019-07-22 Hongsheng Dai , Murray Pollock , Gareth Roberts

We provide a novel characterization of semiparametric efficiency in a generic supervised learning setting where the outcome mean function -- defined as the conditional expectation of the outcome of interest given the other observed…

统计方法学 · 统计学 2025-04-22 Harrison H. Li

In the age of large and heterogeneous datasets, the integration of information from diverse sources is essential to improve parameter estimation. Multi-task learning offers a powerful approach by enabling simultaneous learning across…

统计方法学 · 统计学 2025-07-11 Sohom Bhattacharya , Yongzhuo Chen , Muxuan Liang

In this paper we give a brief review of semiparametric theory, using as a running example the common problem of estimating an average causal effect. Semiparametric models allow at least part of the data-generating process to be unspecified…

统计方法学 · 统计学 2017-09-20 Edward H. Kennedy

In the era of big data, the explosive growth of multi-source heterogeneous data offers many exciting challenges and opportunities for improving the inference of conditional average treatment effects. In this paper, we investigate…

机器学习 · 统计学 2022-11-02 Xinyu Li , Yilin Li , Qing Cui , Longfei Li , Jun Zhou

Recent years have experienced increasing utilization of complex machine learning models across multiple sources of data to inform more generalizable decision-making. However, distribution shifts across data sources and privacy concerns…

统计方法学 · 统计学 2024-05-16 Yi Liu , Alexander W. Levis , Sharon-Lise Normand , Larry Han

We provide finite-sample distribution approximations, that are uniform in the parameter, for inference in linear mixed models. Focus is on variances and covariances of random effects in cases where existing theory fails because their…

统计理论 · 数学 2025-07-29 Karl Oskar Ekvall , Matteo Bottai

Difference-in-differences (DiD) is a cornerstone of causal inference, yet extending it to functional outcomes is not a routine scalar generalization; rather, it entails three fundamental challenges in identification, inference, and…

统计方法学 · 统计学 2026-05-29 Junzhu Nie , Chengxiu Ling , Mengfei Ran

An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable…

We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…

统计理论 · 数学 2018-05-22 Takumi Saegusa

In the analysis of cluster data, the regression coefficients are frequently assumed to be the same across all clusters. This hampers the ability to study the varying impacts of factors on each cluster. In this paper, a semiparametric model…

统计理论 · 数学 2009-08-25 Wenyang Zhang , Jianqing Fan , Yan Sun
‹ 上一页 1 2 3 10 下一页 ›