English
Related papers

Related papers: On the testability of the CAR assumption

200 papers

Distribution shifts introduce uncertainty that undermines the robustness and generalization capabilities of machine learning models. While conventional wisdom suggests that learning causal-invariant representations enhances robustness to…

Machine Learning · Computer Science 2025-05-28 Abbavaram Gowtham Reddy , Celia Rubio-Madrigal , Rebekka Burkholz , Krikamol Muandet

The tetrad constraint is widely used to test whether four observed variables are conditionally independent given a latent variable, based on the fact that if four observed variables following a linear model are mutually independent after…

Methodology · Statistics 2026-04-01 Naiwen Ying , Ping Zhang , Shanshan Luo , Wang Miao

The problem of corrupted data, missing features, or missing modalities continues to plague the modern machine learning landscape. To address this issue, a class of regularization methods that enforce consistency between imputed and fully…

Machine Learning · Computer Science 2026-02-03 Yinsong Wang , Shahin Shahrampour

The conditional randomization test (CRT) was recently proposed to test whether two random variables X and Y are conditionally independent given random variables Z. The CRT assumes that the conditional distribution of X given Z is known…

Machine Learning · Computer Science 2023-04-11 Shuai Li , Ziqi Chen , Hongtu Zhu , Christina Dan Wang , Wang Wen

The statistical censoring setup is extended to the situation when random measures can be assigned to the realization of datapoints, leading to a new way of incorporating expert information into the usual parametric estimation procedures.…

Methodology · Statistics 2023-12-05 Hansjörg Albrecher , Martin Bladt

In this paper we introduce the idea of partially sorting data to design nonparametric tests. This approach gives rise to tests that are sensitive to both the order and the underlying distribution of the data. We focus in particular on a…

Statistics Theory · Mathematics 2022-10-27 Krzysztof Bisewski , H. M. Jansen , Yoni Nazarathy

Missing data problems arise in many applied research studies. They may jeopardize statistical inference of the model of interest, if the missing mechanism is nonignorable, that is, the missing mechanism depends on the missing values…

Statistics Theory · Mathematics 2015-09-15 Wang Miao , Peng Ding , Zhi Geng

We propose a general new method, the conditional permutation test, for testing the conditional independence of variables $X$ and $Y$ given a potentially high-dimensional random vector $Z$ that may contain confounding factors. The proposed…

Methodology · Statistics 2019-05-08 Thomas B. Berrett , Yi Wang , Rina Foygel Barber , Richard J. Samworth

We propose a method to distinguish causal influence from hidden confounding in the following scenario: given a target variable Y, potential causal drivers X, and a large number of background features, we propose a novel criterion for…

Machine Learning · Statistics 2022-02-07 You-Lin Chen , Lenon Minorics , Dominik Janzing

In the analysis of time-to-event data with multiple causes using a competing risks Cox model, often the cause of failure is unknown for some of the cases. The probability of a missing cause is typically assumed to be independent of the…

Methodology · Statistics 2016-08-01 Daniel Nevo , Reiko Nishihara , Shuji Ogino , Molin Wang

The usual parametric models for survival data are of the following form. Some parametrically specified hazard rate $\alpha(s,\theta)$ is assumed for possibly censored random life times $X_1^0,\ldots,X_n^0$; one observes only…

Methodology · Statistics 2026-03-25 Nils Lid Hjort

With nonignorable missing data, likelihood-based inference should be based on the joint distribution of the study variables and their missingness indicators. These joint models cannot be estimated from the data alone, thus requiring the…

Statistics Theory · Mathematics 2017-01-06 Mauricio Sadinle , Jerome P. Reiter

Not all experiments publish their results with a description of the correlations between the data points. This makes it difficult to do hypothesis tests or model fits with that data, since just assuming no correlation can lead to an over-…

Data Analysis, Statistics and Probability · Physics 2021-06-30 Lukas Koch

We present a general principle for estimating a regression function nonparametrically, allowing for a wide variety of data filtering, for example, repeated left truncation and right censoring. Both the mean and the median regression cases…

Statistics Theory · Mathematics 2011-02-10 Oliver Linton , Enno Mammen , Jens Perch Nielsen , Ingrid Van Keilegom

We consider the conditional randomization test as a way to account for covariate imbalance in randomized experiments. The test accounts for covariate imbalance by comparing the observed test statistic to the null distribution of the test…

In conditional copula models, the copula parameter is deterministically linked to a covariate via the calibration function. The latter is of central interest for inference and is usually estimated nonparametrically. However, when a…

Methodology · Statistics 2014-03-19 Elif F. Acar , Radu V. Craiu , Fang Yao

Randomized controlled trials are not only the golden standard in medicine and vaccine trials but have spread to many other disciplines like behavioral economics, making it an important interdisciplinary tool for scientists. When designing…

Methodology · Statistics 2021-11-30 Tassilo Schwarz

We consider the problem of distribution-free predictive inference, with the goal of producing predictive coverage guarantees that hold conditionally rather than marginally. Existing methods such as conformal prediction offer marginal…

Statistics Theory · Mathematics 2020-04-16 Rina Foygel Barber , Emmanuel J. Candès , Aaditya Ramdas , Ryan J. Tibshirani

Invariant Causal Prediction (Peters et al., 2016) is a technique for out-of-distribution generalization which assumes that some aspects of the data distribution vary across the training set but that the underlying causal mechanisms remain…

Machine Learning · Computer Science 2021-03-30 Elan Rosenfeld , Pradeep Ravikumar , Andrej Risteski

A general structural equation model is fitted on a panel data set that consists of $I$ correlated samples. The correlated samples could be data from correlated populations or correlated observations from occasions of panel data. We consider…

Statistics Theory · Mathematics 2007-06-13 Savas Papadopoulos , Yasuo Amemiya