中文
相关论文

相关论文: When can Multi-Site Datasets be Pooled for Regress…

200 篇论文

Probability forecasting is common in the geosciences, the finance sector, and elsewhere. It is sometimes the case that one has multiple probability-forecasts for the same target. How is the information in these multiple forecast systems…

统计方法学 · 统计学 2016-03-02 Sarah Higgins , Hailiang Du , Leonard A. Smith

Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target…

统计方法学 · 统计学 2025-01-09 Lyndon Ang , Robert Clark , Bronwyn Loong , Anders Holmberg

This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed across algorithmic tasks, architectures and optimizers and…

机器学习 · 计算机科学 2026-05-21 Jingwen Liu , Ezra Edelman , Surbhi Goel , Bingbin Liu

In multicenter biomedical research, integrating data from multiple decentralized sites provides more robust and generalizable findings due to its larger sample size and the ability to account for the between-site heterogeneity. However,…

统计方法学 · 统计学 2025-12-29 Xiaokang Liu , Yuchen Yang , Yifei Sun , Jiang Bian , Yanyuan Ma , Raymond J. Carroll , Yong Chen

Causal inference (CI) in observational studies has received a lot of attention in healthcare, education, ad attribution, policy evaluation, etc. Confounding is a typical hazard, where the context affects both, the treatment assignment and…

统计方法学 · 统计学 2021-08-17 Ankit Sharma , Garima Gupta , Ranjitha Prasad , Arnab Chatterjee , Lovekesh Vig , Gautam Shroff

Most machine learning models for predicting clinical outcomes are developed using historical data. Yet, even if these models are deployed in the near future, dataset shift over time may result in less than ideal performance. To capture this…

机器学习 · 计算机科学 2023-06-21 Christina X Ji , Ahmed M Alaa , David Sontag

Data pooling offers various advantages, such as increasing the sample size, improving generalization, reducing sampling bias, and addressing data sparsity and quality, but it is not straightforward and may even be counterproductive.…

计算机视觉与模式识别 · 计算机科学 2024-05-09 Stefan Becker , Jens Bayer , Ronny Hug , Wolfgang Hübner , Michael Arens

As large and powerful neural language models are developed, researchers have been increasingly interested in developing diagnostic tools to probe them. There are many papers with conclusions of the form "observation X is found in model Y",…

计算与语言 · 计算机科学 2022-02-28 Zining Zhu , Jixuan Wang , Bai Li , Frank Rudzicz

Simulation studies are commonly used to evaluate the performance of newly developed meta-analysis methods. For methodology that is developed for an aggregated data meta-analysis, researchers often resort to simulation of the aggregated data…

应用统计 · 统计学 2022-01-19 Edwin R. van den Heuvel , Osama Almalik , Zhuozhao Zhan

Large Language Models (LLMs) are increasingly deployed for clinical reasoning tasks, which inherently require eliciting calibrated probabilistic beliefs based on available evidence. However, real-world clinical data are frequently…

人工智能 · 计算机科学 2026-03-19 Yuta Kobayashi , Vincent Jeanselme , Shalmali Joshi

Distribution shifts remain a fundamental problem for the safe application of machine learning systems. If undetected, they may impact the real-world performance of such systems or will at least render original performance claims invalid. In…

机器学习 · 计算机科学 2023-03-10 Lisa M. Koch , Christian M. Schürch , Christian F. Baumgartner , Arthur Gretton , Philipp Berens

In the problem of composite hypothesis testing, identifying the potential uniformly most powerful (UMP) unbiased test is of great interest. Beyond typical hypothesis settings with exponential family, it is usually challenging to prove the…

统计方法学 · 统计学 2022-08-03 Tianyu Zhan , Jian Kang

In high-stakes domains like healthcare, users often expect that sharing personal information with machine learning systems will yield tangible benefits, such as more accurate diagnoses and clearer explanations of contributing factors.…

机器学习 · 计算机科学 2026-03-18 Louisa Cornelis , Guillermo Bernárdez , Haewon Jeong , Nina Miolane

Large-scale testing is crucial in pandemic containment, but resources are often prohibitively constrained. We study the optimal application of pooled testing for populations that are heterogeneous with respect to an individual's infection…

计算机科学与博弈论 · 计算机科学 2023-09-22 Simon Finster , Michelle González Amador , Edwin Lock , Francisco Marmolejo-Cossío , Evi Micha , Ariel D. Procaccia

Conjoint analysis is a popular experimental design used to measure multidimensional preferences. Researchers examine how varying a factor of interest, while controlling for other relevant factors, influences decision-making. Currently,…

统计方法学 · 统计学 2024-11-20 Dae Woong Ham , Kosuke Imai , Lucas Janson

Statisticians increasingly face the problem to reconsider the adaptability of classical inference techniques. In particular, divers types of high-dimensional data structures are observed in various research areas; disclosing the boundaries…

统计理论 · 数学 2017-06-09 Paavo Sattler , Markus Pauly

Causal analyses for observational studies are often complicated by covariate imbalances among treatment groups, and matching methodologies alleviate this complication by finding subsets of treatment groups that exhibit covariate balance. It…

统计方法学 · 统计学 2021-04-26 Zach Branson

Large-scale behavioral datasets enable researchers to use complex machine learning algorithms to better predict human behavior, yet this increased predictive power does not always lead to a better understanding of the behavior in question.…

计算机与社会 · 计算机科学 2019-05-14 Mayank Agrawal , Joshua C. Peterson , Thomas L. Griffiths

Complex multilayer network datasets have become ubiquitous in various applications, including neuroscience, social sciences, economics, and genetics. Notable examples include brain connectivity networks collected across multiple patients or…

社会与信息网络 · 计算机科学 2026-03-06 Alexander Kagan , Peter W. MacDonald , Elizaveta Levina , Ji Zhu

The problem of multiple hypothesis testing arises when there are more than one hypothesis to be tested simultaneously for statistical significance. This is a very common situation in many data mining applications. For instance, assessing…

机器学习 · 统计学 2009-06-30 Sami Hanhijärvi , Kai Puolamäki , Gemma C. Garriga