中文
相关论文

相关论文: When can Multi-Site Datasets be Pooled for Regress…

200 篇论文

Determining the best model or models for a particular data set, a process known as Bayesian model comparison, is a critical part of probabilistic inference. Typically, this process assumes a fixed model-space (that is, a fixed set of…

定量方法 · 定量生物学 2019-01-08 Thomas HB FitzGerald , Dorothea Hammerer , Thomas D Sambrook , Will D Penny

In identifying infected patients in a population, group testing is an effective method to reduce the number of tests and correct the test errors. In the group testing procedure, tests are performed on pools of specimens collected from…

机器学习 · 统计学 2021-02-10 Ayaka Sakata

It seems as though progressively more people are in the race to upload content, data, and information online; and hospitals haven't neglected this trend either. Hospitals are now at the forefront for multi-site medical data sharing to…

机器学习 · 计算机科学 2022-02-23 Yoo Jeong Ha , Gusang Lee , Minjae Yoo , Soyi Jung , Seehwan Yoo , Joongheon Kim

We show that pooling countries across a panel dimension to macroeconomic data can improve by a statistically significant margin the generalization ability of structural, reduced form, and machine learning (ML) methods to produce…

综合经济学 · 经济学 2022-03-15 Cameron Fen , Samir Undavia

Neuroimaging consortia can enhance reliability and generalizability of findings by pooling data across studies to achieve larger sample sizes. To adjust for site and MRI protocol effects, imaging datasets are often harmonized based on…

Investigators often use multi-source data (e.g., multi-center trials, meta-analyses of randomized trials, pooled analyses of observational cohorts) to learn about the effects of interventions in subgroups of some well-defined target…

统计方法学 · 统计学 2024-02-06 Guanbo Wang , Alexander Levis , Jon Steingrimsson , Issa Dahabreh

This paper presents methods for meta-analysis of $2 \times 2$ tables, both with and without allowing heterogeneity in the treatment effects. Meta-analysis is common in medical research, but most existing methods are unsuited for $2 \times…

统计方法学 · 统计学 2026-02-23 Céline Cunen , Nils Lid Hjort

The identification of the dependent components in multiple data sets is a fundamental problem in many practical applications. The challenge in these applications is that often the data sets are high-dimensional with few observations or…

统计方法学 · 统计学 2023-06-02 Martin Gölz , Tanuj Hasija , Michael Muma , Abdelhak M. Zoubir

The usual problem for group testing is this: For a given number of individuals and a given prevalence, how many tests T* are required to find every infected individual? In real life, however, the problem is usually different: For a given…

应用统计 · 统计学 2021-07-21 Matthew Aldridge

Scholars frequently use covariate balance tests to test the validity of natural experiments and related designs. Unfortunately, when measured covariates are unrelated to potential outcomes, balance is uninformative about key identification…

统计方法学 · 统计学 2025-10-15 Clara Bicalho , Adam Bouyamourn , Thad Dunning

Safely deploying machine learning models to the real world is often a challenging process. Models trained with data obtained from a specific geographic location tend to fail when queried with data obtained elsewhere, agents trained in a…

机器学习 · 计算机科学 2021-11-02 Marco Federici , Ryota Tomioka , Patrick Forré

Data with multiple functional recordings at each observational unit are increasingly common in various fields including medical imaging and environmental sciences. To conduct inference for such observations, we develop a paired two-sample…

统计方法学 · 统计学 2025-06-16 Colin Decker , Dehan Kong , Stanislav Volgushev

Deep neural networks has been increasingly applied in fault diagnostics, where it uses historical data to capture systems behavior, bypassing the need for high-fidelity physical models. However, despite their competence in prediction tasks,…

机器学习 · 计算机科学 2025-09-24 Arman Mohammadi , Mattias Krysander , Daniel Jung , Erik Frisk

Controlling infectious diseases is a major health priority because they can spread and infect humans, thus evolving into epidemics or pandemics. Therefore, early detection of infectious diseases is a significant need, and many researchers…

机器学习 · 计算机科学 2022-06-16 Eman Yahia Alqaissi , Fahd Saleh Alotaibi , Muhammad Sher Ramzan

Imbalanced problems can arise in different real-world situations, and to address this, certain strategies in the form of resampling or balancing algorithms are proposed. This issue has largely been studied in the context of classification,…

机器学习 · 计算机科学 2025-07-17 Juscimara G. Avelino , George D. C. Cavalcanti , Rafael M. O. Cruz

Healthcare datasets often contain groups of highly correlated features, such as features from the same biological system. When feature selection is applied to these datasets to identify the most important features, the biases inherent in…

机器学习 · 计算机科学 2022-07-07 Annette Spooner , Gelareh Mohammadi , Perminder S. Sachdev , Henry Brodaty , Arcot Sowmya

We present a general framework for hypothesis testing on distributions of sets of individual examples. Sets may represent many common data sources such as groups of observations in time series, collections of words in text or a batch of…

统计方法学 · 统计学 2021-02-03 Alexis Bellot , Mihaela van der Schaar

Current statistical inference problems in areas like astronomy, genomics, and marketing routinely involve the simultaneous testing of thousands -- even millions -- of null hypotheses. For high-dimensional multivariate distributions, these…

统计方法学 · 统计学 2017-04-25 Weixin Cai , Nima S. Hejazi , Alan E. Hubbard

A number of applications require two-sample testing on ranked preference data. For instance, in crowdsourcing, there is a long-standing question of whether pairwise comparison data provided by people is distributed similar to…

机器学习 · 统计学 2020-11-20 Charvi Rastogi , Sivaraman Balakrishnan , Nihar B. Shah , Aarti Singh

With an increase of dataset availability, the potential for learning from a variety of data sources has increased. One particular method to improve learning from multiple data sources is to embed the data source during training. This allows…

计算与语言 · 计算机科学 2021-12-08 Rob van der Goot , Miryam de Lhoneux