English
Related papers

Related papers: Model-Free Conditional Feature Screening with Expo…

200 papers

Causal discovery from data affected by unobserved variables is an important but difficult problem to solve. The effects that unobserved variables have on the relationships between observed variables are more complex in nonlinear cases than…

Machine Learning · Computer Science 2021-06-07 Takashi Nicholas Maeda , Shohei Shimizu

This article considers a novel and widely applicable approach to modeling high-dimensional dependent data when a large number of explanatory variables are available and the signal-to-noise ratio is low. We postulate that a $p$-dimensional…

Methodology · Statistics 2024-12-09 Zhaoxing Gao , Ruey S. Tsay

Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the…

Machine Learning · Statistics 2016-01-21 Simone Romano , Nguyen Xuan Vinh , James Bailey , Karin Verspoor

One of the main topics of extreme value analysis is to estimate the extreme value index, an important parameter that controls the tail behavior of the distribution. In many cases, estimating the extreme value index of the target variable…

Methodology · Statistics 2024-10-22 Takuma Yoshida , Yuta Umezu

Screening before model building is a reasonable strategy to reduce the dimension of regression problems. Sure independence screening is an efficient approach to this purpose. It applies the slope estimate of a simple linear regression as a…

Applications · Statistics 2014-01-21 Sheng-Mao Chang

We propose and study properties of maximum likelihood estimators in the class of conditional transformation models. Based on a suitable explicit parameterisation of the unconditional or conditional transformation function, we establish a…

Methodology · Statistics 2019-10-22 Torsten Hothorn , Lisa Möst , Peter Bühlmann

We develop a novel method for counterfactual analysis based on observational data using prediction intervals for units under different exposures. Unlike methods that target heterogeneous or conditional average treatment effects of an…

Statistics Theory · Mathematics 2018-07-18 Dave Zachariah , Petre Stoica

For exchangeable data, mixture models are an extremely useful tool for density estimation due to their attractive balance between smoothness and flexibility. When additional covariate information is present, mixture models can be extended…

Methodology · Statistics 2023-08-01 Sara Wade , Vanda Inacio , Sonia Petrone

In many applications, it is of interest to identify a parsimonious set of features, or panel, from multiple candidates that achieves a desired level of performance in predicting a response. This task is often complicated in practice by…

Methodology · Statistics 2025-10-23 B. D. Williamson , Y. Huang

In the context of the usual calibration model, we consider the case in which the independent variable is unobservable, but a pre-fixed value on its surrogate is available. Thus, considering controlled variables and assuming that the…

Applications · Statistics 2008-02-06 Betsabé G. Blas Achic , Mônica C. Sandoval , Olga Satomi Yoshida

Decision making or scientific discovery pipelines such as job hiring and drug discovery often involve multiple stages: before any resource-intensive step, there is often an initial screening that uses predictions from a machine learning…

Methodology · Statistics 2023-05-30 Ying Jin , Emmanuel J. Candès

When performing regression or classification, we are interested in the conditional probability distribution for an outcome or class variable Y given a set of explanatoryor input variables X. We consider Bayesian models for this task. In…

Machine Learning · Computer Science 2013-02-08 David Heckerman , Christopher Meek

The use of machine learning (ML) in high-stakes societal decisions has encouraged the consideration of fairness throughout the ML lifecycle. Although data integration is one of the primary steps to generate high quality training data, most…

Machine Learning · Computer Science 2022-04-01 Sainyam Galhotra , Karthikeyan Shanmugam , Prasanna Sattigeri , Kush R. Varshney

Certain causal models involving unmeasured variables induce no independence constraints among the observed variables but imply, nevertheless, inequality contraints on the observed distribution. This paper derives a general formula for such…

Artificial Intelligence · Computer Science 2013-02-21 Judea Pearl

Propensity scores are commonly used to reduce the confounding bias in non-randomized observational studies for estimating the average treatment effect. An important assumption underlying this approach is that all confounders that are…

Methodology · Statistics 2022-08-02 Youfei Yu , Jiacong Du , Min Zhang , Zhenke Wu , Andrew M. Ryan , Bhramar Mukherjee

We formulate a causal extension to the recently introduced paradigm of instance-wise feature selection to explain black-box visual classifiers. Our method selects a subset of input features that has the greatest causal effect on the models…

Machine Learning · Computer Science 2021-04-27 Pranoy Panda , Sai Srinivas Kancheti , Vineeth N Balasubramanian

There has been a recent surge in research on causal panel data models, leading to many new estimators for average causal effects. However, researchers have paid less attention to quantifying the precision of these estimators. This paper…

Econometrics · Economics 2025-11-25 Alexander Almeida , Susan Athey , Guido Imbens , Eva Lestant , Alexia Olaizola

This paper studies nonparametric series estimation and inference for the effect of a single variable of interest x on an outcome y in the presence of potentially high-dimensional conditioning variables z. The context is an additively…

Statistics Theory · Mathematics 2020-04-07 Damian Kozbur

Random effects are a flexible addition to statistical models to capture structural heterogeneity in the data, such as spatial dependencies, individual differences, temporal dependencies, or non-linear effects. Testing for the presence (or…

Methodology · Statistics 2024-10-21 Fabio Vieira , Hongwei Zhao , Joris Mulder

We propose a test of the significance of a variable appearing on the Lasso path and use it in a procedure for selecting one of the models of the Lasso path, controlling the Family-Wise Error Rate. Our null hypothesis depends on a set A of…

Methodology · Statistics 2024-09-05 Matthieu Pluntz , Cyril Dalmasso , Pascale Tubert-Bitter , Ismail Ahmed