中文
相关论文

相关论文: When can Multi-Site Datasets be Pooled for Regress…

200 篇论文

It is common when using cross-section or panel data to assign each observation to a cluster and allow for arbitrary patterns of heteroskedasticity and correlation within clusters. For regression models, there are many ways to make…

计量经济学 · 经济学 2026-04-03 James G. MacKinnon

Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is…

机器学习 · 计算机科学 2026-02-10 Ayush Roy , Rudrasis Chakraborty , Lav Varshney , Vishnu Suresh Lokhande

We study adaptive pooling under predictive heterogeneity in high-dimensional multivariate time series forecasting, where global models improve statistical efficiency but may fail to capture heterogeneous predictive structure, while naive…

统计方法学 · 统计学 2026-04-16 Ziling Ma , Ángel López Oriona , Hernando Ombao , Ying Sun

In this paper we discuss how to evaluate the differences between fitted logistic regression models across sub-populations. Our motivating example is in studying computerized diagnosis for learning disabilities, where sub-populations based…

统计方法学 · 统计学 2023-03-24 Guy Ashiri-Prossner , Yuval Benjamini

Collaboration between different data centers is often challenged by heterogeneity across sites. To account for the heterogeneity, the state-of-the-art method is to re-weight the covariate distributions in each site to match the distribution…

机器学习 · 统计学 2024-04-25 Tianyu Guo , Sai Praneeth Karimireddy , Michael I. Jordan

This paper proposes a novel testing procedure for selecting a sparse set of covariates that explains a large dimensional panel. Our selection method provides correct false detection control while having higher power than existing…

计量经济学 · 经济学 2023-03-09 Markus Pelger , Jiacheng Zou

Combining distributions is an important issue in decision theory and Bayesian inference. Logarithmic pooling is a popular method to aggregate expert opinions by using a set of weights that reflect the reliability of each information source.…

Valid statistical inference is challenging when the sample is subject to unknown selection bias. Data integration can be used to correct for selection bias when we have a parallel probability sample from the same population with some common…

统计方法学 · 统计学 2023-07-24 Zhonglei Wang , Shu Yang , Jae Kwang Kim

This work addresses the problem of sharing partial information within social learning strategies. In traditional social learning, agents solve a distributed multiple hypothesis testing problem by performing two operations at each instant:…

信号处理 · 电气工程与系统科学 2022-12-07 Virginia Bordignon , Vincenzo Matta , Ali H. Sayed

In many machine learning for healthcare tasks, standard datasets are constructed by amassing data across many, often fundamentally dissimilar, sources. But when does adding more data help, and when does it hinder progress on desired model…

机器学习 · 计算机科学 2024-08-09 Judy Hanwen Shen , Inioluwa Deborah Raji , Irene Y. Chen

Clinical machine learning applications are often plagued with confounders that can impact the generalizability and predictive performance of the learners. Confounding is especially problematic in remote digital health studies where the…

We develop new algorithms for estimating heterogeneous treatment effects, combining recent developments in transfer learning for neural networks with insights from the causal inference literature. By taking advantage of transfer learning,…

Federated or multi-site studies have distinct advantages over single-site studies, including increased generalizability, the ability to study underrepresented populations, and the opportunity to study rare exposures and outcomes. However,…

机器学习 · 统计学 2023-09-25 Larry Han , Zhu Shen , Jose Zubizarreta

Cross-study replicability is a powerful model evaluation criterion that emphasizes generalizability of predictions. When training cross-study replicable prediction models, it is critical to decide between merging and treating the studies…

机器学习 · 统计学 2022-07-14 Cathy Shyr , Pragya Sur , Giovanni Parmigiani , Prasad Patil

Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the…

计算与语言 · 计算机科学 2024-03-27 Sai Koneru , Jian Wu , Sarah Rajtmajer

Group testing, a problem with diverse applications across multiple disciplines, traditionally assumes independence across nodes' states. Recent research, however, focuses on real-world scenarios that often involve correlations among nodes,…

信息论 · 计算机科学 2025-04-02 Hesam Nikpey , Saswati Sarkar , Shirin Saeedi Bidokhti

A common problem in health research is that we have a large database with many variables measured on a large number of individuals. We are interested in measuring additional variables on a subsample; these measurements may be newly…

统计方法学 · 统计学 2022-03-22 Thomas Lumley , Tong Chen

One of the core applications of machine learning to knowledge discovery consists on building a function (a hypothesis) from a given amount of data (for instance a decision tree or a neural network) such that we can use it afterwards to…

机器学习 · 计算机科学 2007-05-23 Carlos Domingo , Ricard Gavalda , Osamu Watanabe

Estimation of the allele frequency at genetic markers is a key ingredient in biological and biomedical research, such as studies of human genetic variation or of the genetic etiology of heritable traits. As genetic data becomes increasingly…

应用统计 · 统计学 2007-12-18 Marc Coram , Hua Tang

Many biomedical experiments are carried out by pooling individual biological samples. However, pooling samples can potentially hide biological variance and give false confidence concerning the data significance. In the context of microarray…

定量方法 · 定量生物学 2008-06-02 Shu-Dong Zhang , Timothy W. Gant