English
Related papers

Related papers: When is $Y_{obs}$ missing and $Y_{mis}$ observed?

200 papers

Simulation studies are commonly used in methodological research for the empirical evaluation of data analysis methods. They generate artificial data sets under specified mechanisms and compare the performance of methods across conditions.…

Methodology · Statistics 2025-07-11 Samuel Pawel , František Bartoš , Björn S. Siepe , Anna Lohmann

This paper addresses a regression problem in which output label values are the results of sensing the magnitude of a phenomenon. A low value of such labels can mean either that the actual magnitude of the phenomenon was low or that the…

Machine Learning · Computer Science 2023-06-01 Takayuki Katsuki , Takayuki Osogami

A fundamental task in statistical learning is quantifying the joint dependence or association between two continuous random variables. We introduce a novel, fully non-parametric measure that assesses the degree of association between…

In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here,…

Machine Learning · Statistics 2024-03-22 Julie Josse , Jacob M. Chen , Nicolas Prost , Erwan Scornet , Gaël Varoquaux

We show that it is common to lose some datapoints for mea-surements scheduled at regular interval on RIPE Atlas. Thetemporal correlation between missing measurements and con-nection events are analyzed, in the pursuit of…

Networking and Internet Architecture · Computer Science 2017-01-05 Wenqin Shao , Jean-Louis Rougier , François Devienne , Mateusz Viste

This note extends conformal e-prediction to cover the case where there is observed confounding between the random object $X$ and its label $Y$. We consider both the case where the observed data is IID and a case where some dependence…

Statistics Theory · Mathematics 2026-03-13 Vladimir Vovk , Ruodu Wang

Identifying causal relationships from observation data is difficult, in large part, due to the presence of hidden common causes. In some cases, where just the right patterns of conditional independence and dependence lie in the data---for…

Artificial Intelligence · Computer Science 2018-01-08 David Heckerman

Observability is the property that enables to distinguish two different locations in $n$-dimensional state space from a reduced number of measured variables, usually just one. In high-dimensional systems it is therefore important to make…

Neurons and Cognition · Quantitative Biology 2019-05-06 Luis A. Aguirre , Leonardo L. Portes , Christophe Letellier

Structural missingness breaks 'just impute and train': values can be undefined by causal or logical constraints, and the mask may depend on observed variables, unobserved variables (MNAR), and other missingness indicators. It simultaneously…

Cross-match spatially clusters and organizes several astronomical point-source measurements from one or more surveys. Ideally, each object would be found in each survey. Unfortunately, the observation conditions and the objects themselves…

Databases · Computer Science 2007-05-23 Jim Gray , Alex Szalay , Tamas Budavari , Robert Lupton , Maria Nieto-Santisteban , Ani Thakar

Given two relations containing multiple measurements - possibly with uncertainties - our objective is to find which sets of attributes from the first have a corresponding set on the second, using exclusively a sample of the data. This…

Databases · Computer Science 2022-07-20 Alejandro Alvarez-Ayllon , Manuel Palomo-Duarte , Juan-Manuel Dodero

We study the problem of solving a linear sensing system when the observations are unlabeled. Specifically we seek a solution to a linear system of equations y = Ax when the order of the observations in the vector y is unknown. Focusing on…

Information Theory · Computer Science 2015-12-02 Jayakrishnan Unnikrishnan , Saeid Haghighatshoar , Martin Vetterli

Missing data are a common problem for both the construction and implementation of a prediction algorithm. Pattern mixture kernel submodels (PMKS) - a series of submodels for every missing data pattern that are fit using only data from that…

Methodology · Statistics 2017-04-27 Sarah Fletcher Mercaldo , Jeffrey D. Blume

The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown…

Machine Learning · Computer Science 2025-03-11 Seongmin Lee , Marcel Böhme

We investigate model based classification with partially labelled training data. In many biostatistical applications, labels are manually assigned by experts, who may leave some observations unlabelled due to class uncertainty. We analyse…

Methodology · Statistics 2019-04-08 Daniel Ahfock , Geoffrey J. McLachlan

We consider the problem of deciding whether a highly incomplete signal lies within a given subspace. This problem, Matched Subspace Detection, is a classical, well-studied problem when the signal is completely observed. High- dimensional…

Information Theory · Computer Science 2011-01-25 Laura Balzano , Bejamin Recht , Robert Nowak

This paper considers the joint distribution of elements of a random sample and an order statistic of the same sample. \ The motivation for this work stems from the important problem in reliability analysis, to estimate the number of…

Statistics Theory · Mathematics 2019-03-04 Ismihan Bairamov

The standard geostatistical problem is to predict the values of a spatially continuous phenomenon, $S(x)$ say, at locations $x$ using data $(y_i,x_i):i=1,..,n$ where $y_i$ is the realization at location $x_i$ of $S(x_i)$, or of a random…

Applications · Statistics 2014-09-12 Emanuele Giorgi , Peter J. Diggle

Feature models are popular in machine learning and they have been recently used to solve many unsupervised learning problems. In these models every observation is endowed with a finite set of features, usually selected from an infinite…

Statistics Theory · Mathematics 2019-02-28 Fadhel Ayed , Marco Battiston , Federico Camerlenghi , Stefano Favaro

We describe a method that infers whether statistical dependences between two observed variables X and Y are due to a "direct" causal link or only due to a connecting causal path that contains an unobserved variable of low complexity, e.g.,…

Machine Learning · Computer Science 2012-02-20 Dominik Janzing , Eleni Sgouritsa , Oliver Stegle , Jonas Peters , Bernhard Schoelkopf