中文
相关论文

相关论文: RDIS: Random Drop Imputation with Self-Training fo…

200 篇论文

A common approach for handling missing values in data analysis pipelines is multiple imputation via software packages such as MICE (Van Buuren and Groothuis-Oudshoorn, 2011) and Amelia (Honaker et al., 2011). These packages typically assume…

统计方法学 · 统计学 2025-07-23 Trung Phung , Kyle Reese , Ilya Shpitser , Rohit Bhattacharya

Although data may be abundant, complete data is less so, due to missing columns or rows. This missingness undermines the performance of downstream data products that either omit incomplete cases or create derived completed data for…

机器学习 · 计算机科学 2020-06-26 Haw-minn Lu , Giancarlo Perrone , José Unpingco

Model-induced distribution shifts (MIDS) occur as previous model outputs pollute new model training sets over generations of models. This is known as model collapse in the case of generative models, and performative prediction or unfairness…

机器学习 · 计算机科学 2024-03-13 Sierra Wyllie , Ilia Shumailov , Nicolas Papernot

Missing data are ubiquitous in real world applications and, if not adequately handled, may lead to the loss of information and biased findings in downstream analysis. Particularly, high-dimensional incomplete data with a moderate sample…

机器学习 · 计算机科学 2022-12-23 Zongyu Dai , Zhiqi Bu , Qi Long

In classification of incomplete pattern, the missing values can either play a crucial role in the class determination, or have only little influence (or eventually none) on the classification results according to the context. We propose a…

人工智能 · 计算机科学 2016-02-09 Zhun-Ga Liu , Quan Pan , Jean Dezert , Arnaud Martin

When training predictive models on data with missing entries, the most widely used and versatile approach is a pipeline technique where we first impute missing entries and then compute predictions. In this paper, we view prediction with…

机器学习 · 计算机科学 2025-02-25 Dimitris Bertsimas , Arthur Delarue , Jean Pauphilet

Due to complex experimental settings, missing values are common in biomedical data. To handle this issue, many methods have been proposed, from ignoring incomplete instances to various data imputation approaches. With the recent rise of…

机器学习 · 计算机科学 2020-05-14 Kristian Miok , Dong Nguyen-Doan , Marko Robnik-Šikonja , Daniela Zaharie

Causal inference is fundamental to empirical scientific discoveries in natural and social sciences; however, in the process of conducting causal inference, data management problems can lead to false discoveries. Two such problems are (i)…

数据库 · 计算机科学 2023-05-16 Brit Youngmann , Michael Cafarella , Babak Salimi , Anna Zeng

While automated driving is often advertised with better-than-human driving performance, this work reviews that it is nearly impossible to provide direct statistical evidence on the system level that this is actually the case. The amount of…

机器学习 · 计算机科学 2021-12-10 Hanno Gottschalk , Matthias Rottmann , Maida Saltagic

Current clinical decision support systems (DSS) are trained and validated on observational data from the target clinic. This is problematic for treatments validated in a randomized clinical trial (RCT), but not yet introduced in any clinic.…

The missing data problem has been broadly studied in the last few decades and has various applications in different areas such as statistics or bioinformatics. Even though many methods have been developed to tackle this challenge, most of…

Data imputation addresses the challenge of imputing missing values in database instances, ensuring consistency with the overall semantics of the dataset. Although several heuristics which rely on statistical methods, and ad-hoc rules have…

人工智能 · 计算机科学 2024-10-22 Jiang Hua , Michael Bewong , Selasi Kwashie , MD Geaur Rahman , Junwei Hu , Xi Guo , Zaiwen Fen

Diffusion models show promising generation capability for a variety of data. Despite their high generation quality, the inference for diffusion models is still time-consuming due to the numerous sampling iterations required. To accelerate…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Kexun Zhang , Xianjun Yang , William Yang Wang , Lei Li

The problem of machine learning with missing values is common in many areas. A simple approach is to first construct a dataset without missing values simply by discarding instances with missing entries or by imputing a fixed value for each…

机器学习 · 统计学 2018-03-02 Hiroyuki Hanada , Toshiyuki Takada , Jun Sakuma , Ichiro Takeuchi

Data from discovery proteomic and phosphoproteomic experiments typically include missing values that correspond to proteins that have not been identified in the analyzed sample. Replacing the missing values with random numbers, a process…

定量方法 · 定量生物学 2019-10-01 Matus Medo , Daniel M. Aebersold , Michaela Medova

Although substantial efforts have been made to mitigate catastrophic forgetting in continual learning, the intrinsic mechanisms are not well understood. In this work, we demonstrate the existence of "pseudo forgetting": the performance…

机器学习 · 计算机科学 2025-06-10 Huashan Sun , Yizhe Yang , Yinghao Li , Jiawei Li , Yang Gao

Missing data is inevitable in longitudinal clinical trials. Conventionally, the missing at random assumption is assumed to handle missingness, which however is unverifiable empirically. Thus, sensitivity analysis is critically important to…

统计方法学 · 统计学 2022-03-18 Siyi Liu , Shu Yang , Yilong Zhang , Guanghan , Liu

Irregularly sampled time series (ISTS) data has irregular temporal intervals between observations and different sampling rates between sequences. ISTS commonly appears in healthcare, economics, and geoscience. Especially in the medical…

机器学习 · 计算机科学 2020-10-27 Chenxi Sun , Shenda Hong , Moxian Song , Hongyan Li

Missing value imputation is a fundamental challenge in machine intelligence, heavily dependent on data completeness. Current imputation methods often handle numerical and categorical attributes independently, overlooking critical…

机器学习 · 计算机科学 2026-01-09 Xiaopeng Luo , Zexi Tan , Zhuowei Wang

Informed down-sampling (IDS) is known to improve performance in symbolic regression when combined with various selection strategies, especially tournament selection. However, recent work found that IDS's gains are not consistent across all…

神经与进化计算 · 计算机科学 2026-01-28 Alina Geiger , Martin Briesch , Dominik Sobania , Franz Rothlauf
‹ 上一页 1 8 9 10 下一页 ›