English
Related papers

Related papers: Towards quantification of incompleteness in the pa…

200 papers

Clustering attempts to partition data instances into several distinctive groups, while the similarities among data belonging to the common partition can be principally reserved. Furthermore, incomplete data frequently occurs in many…

Machine Learning · Computer Science 2022-08-30 Miao Cheng , Xinge You

This paper contributes a set of quality metrics for identification and visual analysis of structured missingness in high-dimensional data. Missing values in data are a frequent challenge in most data generating domains and may cause a range…

Graphics · Computer Science 2025-05-30 Sara Johansson Fernstad , Sarah Alsufyani , Silvia Del Din , Alison Yarnall , Lynn Rochester

The challenge of CPU evaluation lies in the fact that user-perceived performance metrics can only be measured on an independently running system consisting of the CPU and other indispensable components, and hence it is difficult to…

Performance · Computer Science 2025-07-16 Chenxi Wang , Lei Wang , Wanling Gao , Fanda Fan , Yuchen Su , Yutong Zhou , Yikang Yang , Jianfeng Zhan

Incomplete data are common in practical applications. Most predictive machine learning models do not handle missing values so they require some preprocessing. Although many algorithms are used for data imputation, we do not understand the…

Machine Learning · Statistics 2020-07-07 Katarzyna Woźnica , Przemysław Biecek

This paper addresses one of the most prevalent problems encountered by political scientists working with difference-in-differences (DID) design: missingness in panel data. A common practice for handling missing data, known as complete case…

Methodology · Statistics 2024-12-02 Sooahn Shin

Noisy matrix completion aims at estimating a low-rank matrix given only partial and corrupted entries. Despite substantial progress in designing efficient estimation algorithms, it remains largely unclear how to assess the uncertainty of…

Machine Learning · Statistics 2019-11-15 Yuxin Chen , Jianqing Fan , Cong Ma , Yuling Yan

Choi and Yuan (2025) propose a novel approach to applying matrix completion to the problem of estimating causal effects in panel data. The key insight is that even in the presence of structured patterns of missing data -- i.e. selection…

Methodology · Statistics 2026-02-26 Eli Ben-Michael , Avi Feller

Data integration is a notoriously difficult and heuristic-driven process, especially when ground-truth data are not readily available. This paper presents a measure of uncertainty by providing maximal and minimal ranges of a query outcome…

Databases · Computer Science 2023-09-12 Deniz Turkcapar , Sanjay Krishnan

In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here,…

Machine Learning · Statistics 2024-03-22 Julie Josse , Jacob M. Chen , Nicolas Prost , Erwan Scornet , Gaël Varoquaux

A method for aggregation of expert estimates in small groups is proposed. The method is based on combinatorial approach to decomposition of pair-wise comparison matrices and to processing of expert data. It also uses the basic principles of…

Optimization and Control · Mathematics 2019-11-14 Vitaliy Tsyganok , Sergii Kadenko , Oleh Andriichuk , Pavlo Roik

We present a method to analyze sensitivity of frequentist inferences to potential nonignorability of the missingness mechanism. Rather than starting from the selection model, as is typical in such analyses, we assume that the missingness…

Methodology · Statistics 2023-02-09 Heng Chen , Daniel F. Heitjan

Iterative imputation is a popular tool to accommodate missing data. While it is widely accepted that valid inferences can be obtained with this technique, these inferences all rely on algorithmic convergence. There is no consensus on how to…

Computation · Statistics 2021-10-25 Hanne Ida Oberman , Stef van Buuren , Gerko Vink

In many applications, researchers seek to identify overlapping entities across multiple data files. Record linkage algorithms facilitate this task, in the absence of unique identifiers. As these algorithms rely on semi-identifying…

Methodology · Statistics 2026-04-24 Gauri Kamat , Roee Gutman

Multileaved comparison methods generalize interleaved comparison methods to provide a scalable approach for comparing ranking systems based on regular user interactions. Such methods enable the increasingly rapid research and development of…

Information Retrieval · Computer Science 2017-11-28 Harrie Oosterhuis , Maarten de Rijke

Pairwise comparisons based on human judgements are an effective method for determining rankings of items or individuals. However, as human biases perpetuate from pairwise comparisons to recovered rankings, they affect algorithmic decision…

Computers and Society · Computer Science 2024-08-26 Georg Ahnert , Antonio Ferrara , Claudia Wagner

There is a great need for robust techniques in data mining and machine learning contexts where many standard techniques such as principal component analysis and linear discriminant analysis are inherently susceptible to outliers.…

Methodology · Statistics 2015-09-28 Garth Tarr , Samuel Müller , Neville C. Weber

In this article, we propose two quantitative methods for calculating weight vectors for incomplete pairwise comparison matrices using reference values. Both procedures are extensions of arithmetic and geometric heuristic estimation (HRE)…

Artificial Intelligence · Computer Science 2026-03-09 Konrad Kułakowski , Anna Kędzior , Jacek Szybowski , Jiri Mazurek

As AI systems develop in complexity it is becoming increasingly hard to ensure non-discrimination on the basis of protected attributes such as gender, age, and race. Many recent methods have been developed for dealing with this issue as…

Machine Learning · Computer Science 2020-04-21 Yair Horesh , Noa Haas , Elhanan Mishraky , Yehezkel S. Resheff , Shir Meir Lador

With the continued digitization of societal processes, we are seeing an explosion in available data. This is referred to as big data. In a research setting, three aspects of the data are often viewed as the main sources of challenges when…

Databases · Computer Science 2022-05-24 Lu Chen , Yunjun Gao , Xuan Song , Zheng Li , Yifan Zhu , Xiaoye Miao , Christian S. Jensen

Machine learning techniques have been developed to learn from complete data. When missing values exist in a dataset, the incomplete data should be preprocessed separately by removing data points with missing values or imputation. In this…

Machine Learning · Computer Science 2020-12-25 Hadi A. Khorshidi , Michael Kirley , Uwe Aickelin