English
Related papers

Related papers: When can Multi-Site Datasets be Pooled for Regress…

200 papers

Traditional methods for matching in causal inference are impractical for high-dimensional datasets. They suffer from the curse of dimensionality: exact matching and coarsened exact matching find exponentially fewer matches as the input…

Machine Learning · Statistics 2026-02-12 Oscar Clivio , Fabian Falck , Brieuc Lehmann , George Deligiannidis , Chris Holmes

Curating a large scale medical imaging dataset for machine learning applications is both time consuming and expensive. Balancing the workload between model development, data collection and annotations is difficult for machine learning…

Artificial Intelligence · Computer Science 2022-06-07 Athanasios Vlontzos , Hadrien Reynaud , Bernhard Kainz

To estimate accurately the parameters of a regression model, the sample size must be large enough relative to the number of possible predictors for the model. In practice, sufficient data is often lacking, which can lead to overfitting of…

Applications · Statistics 2024-09-25 Marianne A Jonker , Hassan Pazira , Anthony CC Coolen

We introduce dataset multiplicity, a way to study how inaccuracies, uncertainty, and social bias in training datasets impact test-time predictions. The dataset multiplicity framework asks a counterfactual question of what the set of…

Machine Learning · Computer Science 2023-04-24 Anna P. Meyer , Aws Albarghouthi , Loris D'Antoni

In this paper, we propose algorithms that leverage a known community structure to make group testing more efficient. We consider a population organized in connected communities: each individual participates in one or more communities, and…

Information Theory · Computer Science 2021-03-18 Pavlos Nikolopoulos , Sundara Rajan Srinivasavaradhan , Tao Guo , Christina Fragouli , Suhas Diggavi

Datasets typically contain inaccuracies due to human error and societal biases, and these inaccuracies can affect the outcomes of models trained on such datasets. We present a technique for certifying whether linear regression models are…

Machine Learning · Computer Science 2022-06-09 Anna P. Meyer , Aws Albarghouthi , Loris D'Antoni

The overwhelming majority of empirical research that uses cluster-robust inference assumes that the clustering structure is known, even though there are often several possible ways in which a dataset could be clustered. We propose two tests…

Econometrics · Economics 2023-03-14 James G. MacKinnon , Morten Ørregaard Nielsen , Matthew D. Webb

Multistate models offer a powerful framework for studying disease processes and can be used to formulate intensity-based and more descriptive marginal regression models. They also represent a natural foundation for the construction of joint…

A common technique to reduce model bias in time-series forecasting is to use an ensemble of predictive models and pool their output into an ensemble forecast. In cases where each predictive model has different biases, however, it is not…

Machine Learning · Computer Science 2023-10-26 Dhruvit Patel , Alexander Wikner

Determining whether an algorithmic decision-making system discriminates against a specific demographic typically involves comparing a single point estimate of a fairness metric against a predefined threshold. This practice is statistically…

Machine Learning · Computer Science 2026-03-20 Antonio Ferrara , Francesco Cozzi , Alan Perotti , André Panisson , Francesco Bonchi

Data sharing barriers are paramount challenges arising from multicenter clinical trials where multiple data sources are stored in a distributed fashion at different local study sites. Merging such data sources into a common data storage for…

Methodology · Statistics 2022-04-05 Mengtong Hu , Xu Shi , Peter X. -K. Song

Large crowdsourced datasets are widely used for training and evaluating neural models on natural language inference (NLI). Despite these efforts, neural models have a hard time capturing logical inferences, including those licensed by…

Computation and Language · Computer Science 2019-04-30 Hitomi Yanaka , Koji Mineshima , Daisuke Bekki , Kentaro Inui , Satoshi Sekine , Lasha Abzianidze , Johan Bos

When developing clinical prediction models, it can be challenging to balance between global models that are valid for all patients and personalized models tailored to individuals or potentially unknown subgroups. To aid such decisions, we…

Statistical hypotheses are translations of scientific hypotheses into statements about one or more distributions, often concerning their centre. Tests that assess statistical hypotheses of centre implicitly assume a specific centre, e.g.,…

Methodology · Statistics 2024-02-21 Ryan Thompson , Catherine S. Forbes , Steven N. MacEachern , Mario Peruggia

In cancer research, clustering techniques are widely used for exploratory analyses and dimensionality reduction, playing a critical role in the identification of novel cancer subtypes, often with direct implications for patient management.…

Probabilistic models analyze data by relying on a set of assumptions. Data that exhibit deviations from these assumptions can undermine inference and prediction quality. Robust models offer protection against mismatch between a model's…

Machine Learning · Statistics 2018-06-20 Yixin Wang , Alp Kucukelbir , David M. Blei

Policy decisions often depend on evidence generated elsewhere. We take a Bayesian decision-theoretic approach to choosing where to experiment to optimize external validity. We frame external validity through a policy lens, developing a…

In this work, we consider hypothesis testing and anomaly detection on datasets where each observation is a weighted network. Examples of such data include brain connectivity networks from fMRI flow data, or word co-occurrence counts for…

Machine Learning · Statistics 2018-09-10 Guilherme Gomes , Vinayak Rao , Jennifer Neville

The availability of multi-omics data has revolutionized the life sciences by creating avenues for integrated system-level approaches. Data integration links the information across datasets to better understand the underlying biological…

Methodology · Statistics 2022-09-01 Said el Bouhaddani , Hae-Won Uh , Geurt Jongbloed , Jeanine Houwing-Duistermaat

Randomized trials are often conducted with separate randomizations across multiple sites such as schools, voting districts, or hospitals. These sites can differ in important ways, including the site's implementation, local conditions, and…

Methodology · Statistics 2018-03-19 Lo-Hua Yuan , Avi Feller , Luke W. Miratrix
‹ Prev 1 4 5 6 7 8 10 Next ›