中文
相关论文

相关论文: How do dataset characteristics affect the performa…

200 篇论文

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

统计方法学 · 统计学 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

Numerous works have noted similarities in how machine learning models represent the world, even across modalities. Although much effort has been devoted to uncovering properties and metrics on which these models align, surprisingly little…

机器学习 · 计算机科学 2025-09-30 Zeyu Michael Li , Hung Anh Vu , Damilola Awofisayo , Emily Wenger

In this paper, we propose a propensity score adapted variable selection procedure to select covariates for inclusion in propensity score models, in order to eliminate confounding bias and improve statistical efficiency in observational…

统计方法学 · 统计学 2021-09-14 Kangjie Zhou , Jinzhu Jia

The property of conformal predictors to guarantee the required accuracy rate makes this framework attractive in various practical applications. However, this property is achieved at a price of reduction in precision. In the case of…

机器学习 · 计算机科学 2021-08-13 Marharyta Aleksandrova , Oleg Chertov

The purpose of this project was to collect and analyse data about the comparability and real-life applicability of published results focusing on Microsoft Windows malware, more specifically the impact of dataset size and testing dataset…

密码学与安全 · 计算机科学 2022-06-14 David Illes

Theoretical guarantees for causal inference using propensity scores are partly based on the scores behaving like conditional probabilities. However, scores between zero and one, especially when outputted by flexible statistical estimators,…

统计方法学 · 统计学 2024-11-12 Rom Gutman , Ehud Karavani , Yishai Shimoni

Propensity score (PS) methods are widely used to estimate treatment effects in non-randomized studies. Variance is typically estimated using sandwich or bootstrap methods, which can either treat the PS as estimated or fixed. The latter is…

统计方法学 · 统计学 2025-11-17 Baoshan Zhang , Sean M. O'Brien , Yuan Wu , Laine E. Thomas

While the inverse probability of treatment weighting (IPTW) is a commonly used approach for treatment comparisons in observational data, the resulting estimates may be subject to bias and excessively large variance when there is lack of…

统计方法学 · 统计学 2024-02-13 Zhiqiang Cao , Lama Ghazi , Claudia Mastrogiacomo , Laura Forastiere , F. Perry Wilson , Fan Li

Data mining and machine learning techniques such as classification and regression trees (CART) represent a promising alternative to conventional logistic regression for propensity score estimation. Whereas incomplete data preclude the…

机器学习 · 统计学 2018-07-26 Bas B. L. Penning de Vries , Maarten van Smeden , Rolf H. H. Groenwold

Control for confounders in observational studies was generally handled through stratification and standardization until the 1960s. Standardization typically reweights the stratum-specific rates so that exposure categories become comparable.…

统计方法学 · 统计学 2015-03-11 Niels Keiding , David Clayton

High-dimensional data can be useful for causal inference by providing many confounders that may bolster the plausibility of the ignorability assumption. Propensity score methods are powerful tools for causal inference, are popular in health…

统计方法学 · 统计学 2017-10-10 Jacob Spertus , Sharon-Lise Normand

Choice of training data distribution greatly influences model behavior. Yet, in large-scale settings, precisely characterizing how changes in training data affects predictions is often difficult due to model training costs. Current practice…

机器学习 · 计算机科学 2025-05-23 Alaa Khaddaj , Logan Engstrom , Aleksander Madry

Nonprobability (convenience) samples are increasingly sought to reduce the estimation variance for one or more population variables of interest that are estimated using a randomized survey (reference) sample by increasing the effective…

U.S. state education agencies mark schools displaying achievement gaps between demographic subgroups as needing improvement. Some schools may have few students in these subgroups, such that average end-of-year test scores only noisily…

统计方法学 · 统计学 2025-12-10 Joshua Wasserman , Michael R. Elliott , Ben B. Hansen

One popular method for dealing with large-scale data sets is sampling. For example, by using the empirical statistical leverage scores as an importance sampling distribution, the method of algorithmic leveraging samples and rescales…

统计方法学 · 统计学 2013-06-25 Ping Ma , Michael W. Mahoney , Bin Yu

Dataset scaling, also known as normalization, is an essential preprocessing step in a machine learning pipeline. It is aimed at adjusting attributes scales in a way that they all vary within the same range. This transformation is known to…

机器学习 · 计算机科学 2022-12-26 Lucas B. V. de Amorim , George D. C. Cavalcanti , Rafael M. O. Cruz

This research aims to examine the usefulness of integrating various feature selection methods with regression algorithms for sleep quality prediction. A publicly accessible sleep quality dataset is used to analyze the effect of different…

机器学习 · 计算机科学 2023-03-07 Sai Rohith Tanuku , Venkat Tummala

Propensity score (PS) weighting methods are often used in non-randomized studies to adjust for confounding and assess treatment effects. The most popular among them, the inverse probability weighting (IPW), assigns weights that are…

统计方法学 · 统计学 2020-11-04 Yunji Zhou , Roland A. Matsouaka , Laine Thomas

This paper proposes new estimators for the propensity score that aim to maximize the covariate distribution balance among different treatment groups. Heuristically, our proposed procedure attempts to estimate a propensity score model by…

计量经济学 · 经济学 2020-04-07 Pedro H. C. Sant'Anna , Xiaojun Song , Qi Xu

Calibration is a vital aspect of the performance of risk prediction models, but research in the context of ordinal outcomes is scarce. This study compared calibration measures for risk models predicting a discrete ordinal outcome, and…

统计方法学 · 统计学 2021-11-19 Michael Edlinger , Maarten van Smeden , Hannes F Alber , Maria Wanitschek , Ben Van Calster