中文
相关论文

相关论文: When can Multi-Site Datasets be Pooled for Regress…

200 篇论文

Pooling multiple neuroimaging datasets across institutions often enables improvements in statistical power when evaluating associations (e.g., between risk factors and disease outcomes) that may otherwise be too weak to detect. When there…

机器学习 · 计算机科学 2022-03-30 Vishnu Suresh Lokhande , Rudrasis Chakraborty , Sathya N. Ravi , Vikas Singh

Small sample sizes are common in many disciplines, which necessitates pooling roughly similar datasets across multiple institutions to study weak but relevant associations between images and disease outcomes. Such data often manifest…

机器学习 · 计算机科学 2024-11-19 Sotirios Panagiotis Chytas , Vishnu Suresh Lokhande , Peiran Li , Vikas Singh

Estimating the prevalence of a disease is necessary for evaluating and mitigating risks of its transmission within or between populations. Estimates that consider how prevalence changes with time provide more information about these risks…

应用统计 · 统计学 2021-11-12 Braden Scherting , Alison Peel , Raina Plowright , Andrew Hoegh

Causal inference analyses often use existing observational data, which in many cases has some clustering of individuals. In this paper we discuss propensity score weighting methods in a multilevel setting where within clusters individuals…

应用统计 · 统计学 2020-12-24 Youjin Lee , Trang Q. Nguyen , Elizabeth A. Stuart

As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis testing, we present…

计算机与社会 · 计算机科学 2025-06-19 Harbin Hong , Sebastian Caldas , Liu Leqi

The desire to train complex machine learning algorithms and to increase the statistical power in association studies drives neuroimaging research to use ever-larger datasets. The most obvious way to increase sample size is by pooling scans…

计算机视觉与模式识别 · 计算机科学 2020-10-29 Christian Wachinger , Anna Rieckmann , Sebastian Pölsterl

We study the problem of multiple hypothesis testing for multidimensional data when inter-correlations are present. The problem of multiple comparisons is common in many applications. When the data is multivariate and correlated, existing…

统计理论 · 数学 2015-06-02 Mahdis Azadbakhsh , Xin Gao , Hanna Jankowski

In clinical settings, we often face the challenge of building prediction models based on small observational data sets. For example, such a data set might be from a medical center in a multi-center study. Differences between centers might…

Empirical claims often rely on one population, design, and analysis. Many-analysts, multiverse, and robustness studies expose how results can vary across plausible analytic choices. Synthesizing these results, however, is nontrivial as all…

统计方法学 · 统计学 2025-11-24 František Bartoš , Suzanne Hoogeveen , Alexandra Sarafoglou , Samuel Pawel

Statistical inference for large data panels is omnipresent in modern economic applications. An important benefit of panel analysis is the possibility to reduce noise and thus to guarantee stable inference by intersectional pooling. However,…

统计方法学 · 统计学 2022-12-15 Tim Kutta , Holger Dette

Due to large number of entities in biomedical knowledge bases, only a small fraction of entities have corresponding labelled training data. This necessitates entity linking models which are able to link mentions of unseen entities using…

计算与语言 · 计算机科学 2021-04-12 Rico Angell , Nicholas Monath , Sunil Mohan , Nishant Yadav , Andrew McCallum

Panels with large time $(T)$ and cross-sectional $(N)$ dimensions are a key data structure in social sciences and other fields. A central question in panel data analysis is whether to pool data across individuals or to estimate separate…

统计方法学 · 统计学 2025-12-18 Tim Kutta , Martin Schumann , Holger Dette

In many mobile health interventions, treatments should only be delivered in a particular context, for example when a user is currently stressed, walking or sedentary. Even in an optimal context, concerns about user burden can restrict which…

机器学习 · 计算机科学 2018-12-04 Sabina Tomkins , Predrag Klasnja , Susan Murphy

Large scale disease screening is a complicated process in which high costs must be balanced against pressing public health needs. When the goal is screening for infectious disease, one approach is group testing in which samples are…

应用统计 · 统计学 2021-03-02 Gregory Haber , Yaakov Malinovsky , Paul S. Albert

Two-sample testing is a fundamental problem in statistics. Despite its long history, there has been renewed interest in this problem with the advent of high-dimensional and complex data. Specifically, in the machine learning literature,…

统计方法学 · 统计学 2019-11-19 Ilmun Kim , Ann B. Lee , Jing Lei

We consider high-dimensional regression over subgroups of observations. Our work is motivated by biomedical problems, where disease subtypes, for example, may differ with respect to underlying regression models, but sample sizes at the…

Algorithms and technologies are essential tools that pervade all aspects of our daily lives. In the last decades, health care research benefited from new computer-based recruiting methods, the use of federated architectures for data…

计算机与社会 · 计算机科学 2023-01-26 Chiara Criscuolo , Tommaso Dolci , Mattia Salnitri

Comparing two population means of network data is of paramount importance in a wide range of scientific applications. Many existing network inference solutions focus on global testing of entire networks, without comparing individual network…

统计方法学 · 统计学 2019-10-10 Yin Xia , Lexin Li

Pooling publicly-available MRI data from multiple sites allows to assemble extensive groups of subjects, increase statistical power, and promote data reuse with machine learning techniques. The harmonization of multicenter data is necessary…

机器学习 · 计算机科学 2024-02-02 Chiara Marzi , Marco Giannelli , Andrea Barucci , Carlo Tessa , Mario Mascalchi , Stefano Diciotti

In recent years, it has become common practice in neuroscience to use networks to summarize relational information in a set of measurements, typically assumed to be reflective of either functional or structural relationships between regions…

应用统计 · 统计学 2017-03-20 Cedric E. Ginestet , Jun Li , Prakash Balachandran , Steven Rosenberg , Eric D. Kolaczyk
‹ 上一页 1 2 3 10 下一页 ›