English
Related papers

Related papers: IMAGIC-500: IMputation benchmark on A Generative I…

200 papers

Colleges and universities use predictive analytics in a variety of ways to increase student success rates. Despite the potential for predictive analytics, two major barriers exist to their adoption in higher education: (a) the lack of…

Computers and Society · Computer Science 2023-01-02 Hadis Anahideh , Parian Haghighat , Nazanin Nezami , Denisa G`andara

Ensuring the reproducibility of scientific work is crucial as it allows the consistent verification of scientific claims and facilitates the advancement of knowledge by providing a reliable foundation for future research. However,…

Software Engineering · Computer Science 2025-04-14 Lázaro Costa , Susana Barbosa , Jácome Cunha

The Synthetic Minority Over-sampling Technique (SMOTE) is one of the most widely used methods for addressing class imbalance and generating synthetic data. Despite its popularity, little attention has been paid to its privacy implications;…

Cryptography and Security · Computer Science 2026-03-03 Georgi Ganev , Reza Nazari , Rees Davison , Amir Dizche , Xinmin Wu , Ralph Abbey , Jorge Silva , Emiliano De Cristofaro

Missing data are often dealt with multiple imputation. A crucial part of the multiple imputation process is selecting sensible models to generate plausible values for incomplete data. A method based on posterior predictive checking is…

Computation · Statistics 2026-05-14 Mingyang Cai , Stef van Buuren , Gerko Vink

Tabular foundation models are pre-trained on one of three classes of corpus: curated datasets drawn from benchmark repositories, tables harvested at scale from the web, or synthetic tables sampled from a parametric generative prior. Despite…

Artificial Intelligence · Computer Science 2026-05-08 Alex O. Davies , Telmo de Menezes e Silva Filho , Nirav Ajmeri

Background: Existing guidelines for handling missing data are generally not consistent with the goals of prediction modelling, where missing data can occur at any stage of the model pipeline. Multiple imputation (MI), often heralded as the…

Methodology · Statistics 2022-06-27 Rose Sisk , Matthew Sperrin , Niels Peek , Maarten van Smeden , Glen P. Martin

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

Methodology · Statistics 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

Dataset replication is a useful tool for assessing whether improvements in test accuracy on a specific benchmark correspond to improvements in models' ability to generalize reliably. In this work, we present unintuitive yet significant ways…

The Ising model has become a popular psychometric model for analyzing item response data. The statistical inference of the Ising model is typically carried out via a pseudo-likelihood, as the standard likelihood approach suffers from a high…

Methodology · Statistics 2025-01-08 Siliang Zhang , Yunxiao Chen

This study investigates the impact of masking strategies on time series imputation models in healthcare settings. While current approaches predominantly rely on random masking for model evaluation, this practice fails to capture the…

Machine Learning · Computer Science 2025-02-05 Linglong Qian , Yiyuan Yang , Wenjie Du , Jun Wang , Richard Dobsoni , Zina Ibrahim

Person re-identification (Re-ID) benefits greatly from the accurate annotations of existing datasets (e.g., CUHK03 [1] and Market-1501 [2]), which are quite expensive because each image in these datasets has to be assigned with a proper…

Computer Vision and Pattern Recognition · Computer Science 2020-07-16 Guangrun Wang , Guangcong Wang , Xujie Zhang , Jianhuang Lai , Zhengtao Yu , Liang Lin

Data-centric technologies provide exciting opportunities, but recent research has shown how lack of representation in datasets, often as a result of systemic inequities and socioeconomic disparities, can produce inequitable outcomes that…

Human-Computer Interaction · Computer Science 2025-01-15 Gabriella Thompson , Ebtesam Al Haque , Paulette Blanc , Meme Styles , Denae Ford , Angela D. R. Smith , Brittany Johnson

Despite the progress made in deepfake detection research, recent studies have shown that biases in the training data for these detectors can result in varying levels of performance across different demographic groups, such as race and…

Machine Learning · Computer Science 2025-01-03 Uzoamaka Ezeakunne , Chrisantus Eze , Xiuwen Liu

Tabular data is one of the most prevalent and important data formats in real-world applications such as healthcare, finance, and education. However, its effective use in machine learning is often constrained by data scarcity, privacy…

Machine Learning · Computer Science 2025-07-18 Ruxue Shi , Yili Wang , Mengnan Du , Xu Shen , Yi Chang , Xin Wang

Survey data collection often is plagued by unit and item nonresponse. To reduce reliance on strong assumptions about the missingness mechanisms, statisticians can use information about population marginal distributions known, for example,…

Methodology · Statistics 2024-06-10 Yanjiao Yang , Jerome P. Reiter

A major limitation to advances in fingerprint spoof detection is the lack of publicly available, large-scale fingerprint spoof datasets, a problem which has been compounded by increased concerns surrounding privacy and security of biometric…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Steven A. Grosz , Anil K. Jain

Datasets with missing values are very common in real world applications. GAIN, a recently proposed deep generative model for missing data imputation, has been proved to outperform many state-of-the-art methods. But GAIN only uses a…

Machine Learning · Computer Science 2021-04-07 Yufeng Wang , Dan Li , Xiang Li , Min Yang

Handling missing data is crucial in machine learning, but many datasets contain gaps due to errors or non-response. Unlike traditional methods such as listwise deletion, which are simple but inadequate, the literature offers more…

Cryptography and Security · Computer Science 2024-05-30 Julia Jentsch , Ali Burak Ünal , Şeyma Selcan Mağara , Mete Akgün

Synthetic data generation is increasingly used in machine learning for training and data augmentation. Yet, current strategies often rely on external foundation models or datasets, whose usage is restricted in many scenarios due to policy…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Parsa Rahimi , Sebastien Marcel

Due to their data-driven nature, Machine Learning (ML) models are susceptible to bias inherited from data, especially in classification problems where class and group imbalances are prevalent. Class imbalance (in the classification target)…

Machine Learning · Computer Science 2024-09-10 Emmanouil Panagiotou , Arjun Roy , Eirini Ntoutsi
‹ Prev 1 4 5 6 7 8 10 Next ›