English
Related papers

Related papers: Inference with Transposable Data: Modeling the Eff…

200 papers

Datasets with missing values are very common on industry applications, and they can have a negative impact on machine learning models. Recent studies introduced solutions to the problem of imputing missing values based on deep generative…

Machine Learning · Computer Science 2019-02-28 Ramiro D. Camino , Christian A. Hammerschmidt , Radu State

High-dimensional group inference is an essential part of statistical methods for analysing complex data sets, including hierarchical testing, tests of interaction, detection of heterogeneous treatment effects and inference for local…

Methodology · Statistics 2020-12-01 Zijian Guo , Claude Renaux , Peter Bühlmann , T. Tony Cai

In addition to the commonly analyzed measures of location, dispersion measurements such as variance and correlation provide many valuable information. Consequently, they play a crucial role in multivariate statistics, which leads to tests…

Computation · Statistics 2025-09-26 Paavo Sattler , Svenja Jedhoff

Imputation methods play a critical role in enhancing the quality of practical time-series data, which often suffer from pervasive missing values. Recently, diffusion-based generative imputation methods have demonstrated remarkable success…

Machine Learning · Computer Science 2025-10-03 Zeqi Ye , Minshuo Chen

Covariance matrices of random vectors contain information that is crucial for modelling. Specific structures and patterns of the covariances (or correlations) may be used to justify parametric models, e.g., autoregressive models. Until now,…

Methodology · Statistics 2025-02-11 Paavo Sattler , Dennis Dobler

This paper studies nonparametric series estimation and inference for the effect of a single variable of interest x on an outcome y in the presence of potentially high-dimensional conditioning variables z. The context is an additively…

Statistics Theory · Mathematics 2020-04-07 Damian Kozbur

Inferring causal effects of a treatment, intervention or policy from observational data is central to many applications. However, state-of-the-art methods for causal inference seldom consider the possibility that covariates have missing…

Methodology · Statistics 2020-02-26 Imke Mayer , Julie Josse , Félix Raimundo , Jean-Philippe Vert

We consider identification, inference and validation of linear panel data models when both factors and factor loadings are accounted for by a nonparametric function. This general specification encompasses rather popular models such as the…

Econometrics · Economics 2025-06-13 Juan M. Rodriguez-Poo , Alexandra Soberon , Stefan Sperlich

Model explainability is crucial for human users to be able to interpret how a proposed classifier assigns labels to data based on its feature values. We study generalized linear models constructed using sets of feature value rules, which…

Machine Learning · Statistics 2023-11-06 Sanjeeb Dash , Soumyadip Ghosh , Joao Goncalves , Mark S. Squillante

Hidden variable graphical models can sometimes imply constraints on the observable distribution that are more complex than simple conditional independence relations. These observable constraints can falsify assumptions of the model that…

Methodology · Statistics 2026-05-12 Michael C. Sachs , Erin E. Gabriel , Robin J. Evans , Arvid Sjölander

This paper clarifies a fundamental difference between causal inference and traditional statistical inference by formalizing a mathematical distinction between their respective parameters. We connect two major approaches to causal inference,…

Methodology · Statistics 2025-08-29 Muye Liu , Jun Xie

In this paper, we develop invariance-based procedures for testing and inference in high-dimensional regression models. These procedures, also known as randomization tests, provide several important advantages. First, for the global null…

Methodology · Statistics 2023-12-27 Wenxuan Guo , Panos Toulis

An important problem in network analysis is predicting a node attribute using both network covariates, such as graph embedding coordinates or local subgraph counts, and conventional node covariates, such as demographic characteristics.…

Methodology · Statistics 2023-02-24 Robert Lunde , Elizaveta Levina , Ji Zhu

Tabular data (or tables) are the most widely used data format in machine learning (ML). However, ML models often assume the table structure keeps fixed in training and testing. Before ML modeling, heavy data cleaning is required to merge…

Machine Learning · Computer Science 2022-09-19 Zifeng Wang , Jimeng Sun

When inferring parameters from a Gaussian-distributed data set by computing a likelihood, a covariance matrix is needed that describes the data errors and their correlations. If the covariance matrix is not known a priori, it may be…

Cosmology and Nongalactic Astrophysics · Physics 2016-01-27 Elena Sellentin , Alan F. Heavens

We consider functional linear regression models where functional outcomes are associated with scalar predictors by coefficient functions with shape constraints, such as monotonicity and convexity, that apply to sub-domains of interest. To…

Methodology · Statistics 2025-05-09 Kyunghee Han , Yeonjoo Park , Soo-Young Kim

We study a model where one target variable Y is correlated with a vector X:=(X_1,...,X_d) of predictor variables being potential causes of Y. We describe a method that infers to what extent the statistical dependences between X and Y are…

Machine Learning · Statistics 2017-10-11 Dominik Janzing , Bernhard Schoelkopf

We propose robust methods for inference on the effect of a treatment variable on a scalar outcome in the presence of very many controls. Our setting is a partially linear model with possibly non-Gaussian and heteroscedastic disturbances.…

Methodology · Statistics 2017-10-05 Alexandre Belloni , Victor Chernozhukov , Christian Hansen

In this study, we propose a projection estimation method for large-dimensional matrix factor models with cross-sectionally spiked eigenvalues. By projecting the observation matrix onto the row or column factor space, we simplify factor…

Methodology · Statistics 2020-12-04 Long Yu , Yong He , Xin-bing Kong , Xinsheng Zhang

Multimodal data, where different types of data are collected from the same subjects, are fast emerging in a large variety of scientific applications. Factor analysis is commonly used in integrative analysis of multimodal data, and is…

Statistics Theory · Mathematics 2021-03-31 Quefeng Li , Lexin Li