English
Related papers

Related papers: Removing batch effects for prediction problems wit…

200 papers

Stochastic inverse problems are generally solved by some form of finite sampling of a space of uncertain parameters. For computationally expensive models, surrogate response surfaces are often employed to increase the number of samples used…

Numerical Analysis · Mathematics 2018-07-04 Steven Mattis , Barbara Wohlmuth

This paper develops new variance-reduction techniques for the forward-reflected-backward splitting (FRBS) method to solve a class of possibly nonmonotone stochastic composite inclusions. Unlike unbiased estimators such as mini-batching,…

Machine Learning · Computer Science 2026-03-17 Quoc Tran-Dinh , Nghia Nguyen-Trung

Black-box variational inference (BBVI) scales poorly to high-dimensional problems when it is used to estimate a multivariate Gaussian approximation with a full covariance matrix. In this paper, we extend the batch-and-match (BaM) framework…

Machine Learning · Statistics 2025-04-03 Chirag Modi , Diana Cai , Lawrence K. Saul

Evolutionary Algorithms (EAs) are often challenging to apply in real-world settings since evolutionary computations involve a large number of evaluations of a typically expensive fitness function. For example, an evaluation could involve…

Neural and Evolutionary Computing · Computer Science 2024-04-08 Mohammed Ghaith Altarabichi , Sławomir Nowaczyk , Sepideh Pashami , Peyman Sheikholharam Mashhadi

Confounding is a significant obstacle to unbiased estimation of causal effects from observational data. For settings with high-dimensional covariates -- such as text data, genomics, or the behavioral social sciences -- researchers have…

Artificial Intelligence · Computer Science 2024-02-01 Katherine A. Keith , Sergey Feldman , David Jurgens , Jonathan Bragg , Rohit Bhattacharya

Fully test-time adaptation (FTTA) adapts a model that is trained on a source domain to a target domain during the testing phase, where the two domains follow different distributions and source data is unavailable during the training phase.…

Artificial Intelligence · Computer Science 2023-12-15 Houcheng Su , Daixian Liu , Mengzhu Wang , Wei Wang

The identification of surrogate markers is motivated by their potential to make decisions sooner about a treatment effect. However, few methods have been developed to actually use a surrogate marker to test for a treatment effect in a…

Methodology · Statistics 2024-09-17 Layla Parast , Jay Bartroff

Inferring the causal effect of a treatment on an outcome in an observational study requires adjusting for observed baseline confounders to avoid bias. However, adjusting for all observed baseline covariates, when only a subset are…

Methodology · Statistics 2021-02-04 Wen Wei Loh , Stijn Vansteelandt

Given the long follow-up periods that are often required for treatment or intervention studies, the potential to use surrogate markers to decrease the required follow-up time is a very attractive goal. However, previous studies have shown…

Methodology · Statistics 2016-08-12 Layla Parast , Tianxi Cai , Lu Tian

In the paper, a multi-objective evolutionary surrogate-assisted approach for the fast and effective generative design of coastal breakwaters is proposed. To approximate the computationally expensive objective functions, the deep…

Neural and Evolutionary Computing · Computer Science 2022-10-28 Nikita O. Starodubcev , Nikolay O. Nikitin , Anna V. Kalyuzhnaya

In genome-wide association studies (GWAS), hundreds of thousands of genetic markers (SNPs) are tested for association with a trait or phenotype. Reported effects tend to be larger in magnitude than the true effects of these markers, the…

Methodology · Statistics 2010-10-25 Michael E. Goddard , Naomi R. Wray , Klara Verbyla , Peter M. Visscher

Anomaly Detection in multivariate time series is a major problem in many fields. Due to their nature, anomalies sparsely occur in real data, thus making the task of anomaly detection a challenging problem for classification algorithms to…

Machine Learning · Computer Science 2023-08-08 Anastasios Iliopoulos , John Violos , Christos Diou , Iraklis Varlamis

Single-cell perturbation modeling is fundamental for understanding and predicting cellular responses to genetic perturbations. However, existing approaches, from causal representation learning to foundation models, often struggle with an…

Machine Learning · Computer Science 2026-05-20 Wenkang Jiang , Yuhang Liu , Yichao Cai , Erdun Gao , Jiayi Dong , Ehsan Abbasnejad , Lina Yao , Javen Qinfeng Shi

In this work we provide a theoretical framework for structured prediction that generalizes the existing theory of surrogate methods for binary and multiclass classification based on estimating conditional probabilities with smooth convex…

Machine Learning · Computer Science 2019-02-14 Alex Nowak-Vila , Francis Bach , Alessandro Rudi

Traditional variable selection methods could fail to be sign consistent when irrepresentable conditions are violated. This is especially critical in high-dimensional settings when the number of predictors exceeds the sample size. In this…

Methodology · Statistics 2022-04-26 Fei Xue , Annie Qu

Data integration methods aim to extract low-dimensional embeddings from high-dimensional outcomes to remove unwanted variations, such as batch effects and unmeasured covariates, across heterogeneous datasets. However, multiple hypothesis…

Methodology · Statistics 2025-12-15 Jin-Hong Du , Kathryn Roeder , Larry Wasserman

Covariate adjustment and methods of incorporating historical data in randomized clinical trials (RCTs) each provide opportunities to increase trial power. We unite these approaches for the analysis of RCTs with binary outcomes based on the…

Methodology · Statistics 2022-12-21 Alyssa M. Vanderbeek , Jessica L. Ross , David P. Miller , Alejandro Schuler

Randomization tests are a popular method for testing causal effects in clinical trials with finite-sample validity. In the presence of heterogeneous treatment effects, it is often of interest to select a subgroup that benefits from the…

Methodology · Statistics 2025-04-29 Zijun Gao

Causal effect estimation is a critical task in statistical learning that aims to find the causal effect on subjects by identifying causal links between a number of predictor (or, explanatory) variables and the outcome of a treatment. In a…

Methodology · Statistics 2024-11-26 Tathagata Basu , Matthias C. M. Troffaes

The estimation of causal treatment effects from observational data is a fundamental problem in causal inference. To avoid bias, the effect estimator must control for all confounders. Hence practitioners often collect data for as many…

Machine Learning · Statistics 2020-11-05 Kristjan Greenewald , Dmitriy Katz-Rogozhnikov , Karthik Shanmugam