English
Related papers

Related papers: Two-phase validation sampling via principal compon…

200 papers

Factor analysis (FA) and principal component analysis (PCA) are popular statistical methods for summarizing and explaining the variability in multivariate datasets. By default, FA and PCA assume the number of components or factors to be…

Methodology · Statistics 2022-05-17 Chetkar Jha , Ian Barnett

The imputation of missing values in multivariate time series (MTS) data is critical in ensuring data quality and producing reliable data-driven predictive models. Apart from many statistical approaches, a few recent studies have proposed…

Machine Learning · Computer Science 2023-05-17 Maksims Kazijevs , Manar D. Samad

Background: Embedded feature selection in high-dimensional data with very small sample sizes requires optimized hyperparameters for the model building process. For this hyperparameter optimization, nested cross-validation must be applied to…

Machine Learning · Computer Science 2022-09-13 Sigrun May , Sven Hartmann , Frank Klawonn

We introduce a method to estimate simultaneously the tail and the threshold parameters of an extreme value regression model. This standard model finds its use in finance to assess the effect of market variables on extreme loss distributions…

Methodology · Statistics 2023-04-17 Julien Hambuckers , Marie Kratz , Antoine Usseglio-Carleve

We introduce a powerful deep classifier two-sample test for high-dimensional data based on E-values, called E-value Classifier Two-Sample Test (E-C2ST). Our test combines ideas from existing work on split likelihood ratio tests and…

Methodology · Statistics 2024-05-01 Teodora Pandeva , Tim Bakker , Christian A. Naesseth , Patrick Forré

Bias in causal comparisons has a direct correspondence with distributional imbalance of covariates between treatment groups. Weighting strategies such as inverse propensity score weighting attempt to mitigate bias by either modeling the…

Methodology · Statistics 2022-03-14 Jared D. Huling , Simon Mak

Randomized controlled trials (RCTs) are widely regarded as the gold standard for causal inference in biomedical research. For instance, when estimating the average treatment effect on the treated (ATT), a doubly robust estimation procedure…

Methodology · Statistics 2025-09-26 Chi-Shian Dai , Chao Ying , Yang Ning , Jiwei Zhao

Agentic Test-Time Scaling (TTS) has delivered state-of-the-art (SOTA) performance on complex software engineering tasks such as code generation and bug fixing. However, its practical adoption remains limited due to significant computational…

Software Engineering · Computer Science 2026-05-14 Chenhui Mao , Yuanting Lei , Zhixiang Wei , Ming Liang , Zhixiang Wang , Jingxuan Xu , Dajun Chen , Wei Jiang , Yong Li

Information extracted from electrohysterography recordings could potentially prove to be an interesting additional source of information to estimate the risk on preterm birth. Recently, a large number of studies have reported near-perfect…

In precision medicine, one of the most important problems is estimating the optimal individualized treatment rules (ITR), which typically involves recommending treatment decisions based on fully observed individual characteristics of…

Methodology · Statistics 2025-10-15 Yue Zhang , Shanshan Luo , Zhi Geng , Yangbo He

In high-dimensional data analysis, bi-level sparsity is often assumed when covariates function group-wisely and sparsity can appear either at the group level or within certain groups. In such cases, an ideal model should be able to…

Methodology · Statistics 2021-09-14 Bin Luo , Xiaoli Gao

We present a new procedure for enhanced variable selection for component-wise gradient boosting. Statistical boosting is a computational approach that emerged from machine learning, which allows to fit regression models in the presence of…

Longitudinal observations are sometimes costly or not available. Cross sectional sampling can be an alternative. Observations are drawn then at a specific point in time from a population of durations whose distributions satisfy a {\em core…

Statistics Theory · Mathematics 2007-06-13 Bert van Es , Chris A. J. Klaassen , Philip J. Mokveld

Deep learning models for pulmonary disease screening from Computed Tomography (CT) scans promise to alleviate the immense workload on radiologists. Still, their high computational cost, stemming from processing entire 3D volumes, remains a…

Image and Video Processing · Electrical Eng. & Systems 2026-03-19 Qian Shao , Bang Du , Yixuan Wu , Zepeng Li , Qiyuan Chen , Qianqian Tang , Jian Wu , Jintai Chen , Hongxia Xu

In the first stage of a two-stage study, the researcher uses a statistical model to impute the unobserved exposures. In the second stage, imputed exposures serve as covariates in epidemiological models. Imputation error in the first stage…

Applications · Statistics 2021-07-19 Ron Sarafian , Itai Kloog , Jonathan D. Rosenblatt

Evolution Strategies (ES) emerged as a scalable alternative to popular Reinforcement Learning (RL) techniques, providing an almost perfect speedup when distributed across hundreds of CPU cores thanks to a reduced communication overhead.…

Machine Learning · Statistics 2018-11-13 Víctor Campos , Xavier Giro-i-Nieto , Jordi Torres

When multiple investigators analyze a common dataset, the data reuse induces dependence across testing procedures, affecting the distribution of errors. Existing techniques of managing dependent tests require either cross-study coordination…

Statistics Theory · Mathematics 2026-04-10 Reid Dale , Jordan Rodu , Maria E. Currie , Mike Baiocchi

Longitudinal imaging studies are essential to understanding the neural development of neuropsychiatric disorders, substance use disorders, and the normal brain. The main objective of this paper is to develop a two-stage adjusted…

Applications · Statistics 2011-08-12 Xiaoyan Shi , Joseph G. Ibrahim , Jeffrey Lieberman , Martin Styner , Yimei Li , Hongtu Zhu

Measurement error arises commonly in clinical research settings that rely on data from electronic health records or large observational cohorts. In particular, self-reported outcomes are typical in cohort studies for chronic diseases such…

Methodology · Statistics 2021-02-08 Lillian A. Boe , Lesley F. Tinker , Pamela A. Shaw

Large observational datasets, including those derived from electronic health records, are a valuable resource for medical research but are often affected by missingness, measurement error, and misclassification. Two-phase sampling with…

Methodology · Statistics 2026-03-23 Jasper B. Yang , Bryan E. Shepherd , Thomas Lumley , Pamela A. Shaw