English
Related papers

Related papers: Data-driven model selection within the matrix comp…

200 papers

Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait.…

Machine Learning · Statistics 2012-05-31 Chamont Wang , Jana Gevertz , Chaur-Chin Chen , Leonardo Auslender

Latent or unobserved phenomena pose a significant difficulty in data analysis as they induce complicated and confounding dependencies among a collection of observed variables. Factor analysis is a prominent multivariate statistical modeling…

Methodology · Statistics 2020-06-22 Armeen Taeb , Venkat Chandrasekaran

Model selection is critical in the modern statistics and machine learning community. However, most existing works do not apply to heavy-tailed data, which are commonly encountered in real applications, such as the single-cell multiomics…

Methodology · Statistics 2023-05-11 Zhanrui Cai

In covariance matrix estimation, one of the challenges lies in finding a suitable model and an efficient estimation method. Two commonly used modelling approaches in the literature involve imposing linear restrictions on the covariance…

Statistics Theory · Mathematics 2024-05-09 Piotr Zwiernik

We present a data-driven algorithm for efficiently computing stochastic control policies for general joint chance constrained optimal control problems. Our approach leverages the theory of kernel distribution embeddings, which allows…

Systems and Control · Electrical Eng. & Systems 2022-02-10 Adam J. Thorpe , Thomas Lew , Meeko M. K. Oishi , Marco Pavone

Multivariate Gaussian is often used as a first approximation to the distribution of high-dimensional data. Determining the parameters of this distribution under various constraints is a widely studied problem in statistics, and is often…

Statistics Theory · Mathematics 2016-02-09 Samuel Balmand , Arnak Dalalyan

The increasing popularity of regression discontinuity methods for causal inference in observational studies has led to a proliferation of different estimating strategies, most of which involve first fitting non-parametric regression models…

Methodology · Statistics 2018-06-11 Guido Imbens , Stefan Wager

Over the last decades, many prognostic models based on artificial intelligence techniques have been used to provide detailed predictions in healthcare. Unfortunately, the real-world observational data used to train and validate these models…

Machine Learning · Computer Science 2023-11-21 Alice Bernasconi , Alessio Zanga , Peter J. F. Lucas , Marco Scutari , Fabio Stella

We propose dimension reduction methods for sparse, high-dimensional multivariate response regression models. Both the number of responses and that of the predictors may exceed the sample size. Sometimes viewed as complementary, predictor…

Statistics Theory · Mathematics 2013-02-14 Florentina Bunea , Yiyuan She , Marten H. Wegkamp

This paper presents how the most recent improvements made on covariance matrix estimation and model order selection can be applied to the portfolio optimisation problem. The particular case of the Maximum Variety Portfolio is treated but…

Applications · Statistics 2018-04-03 Emmanuelle Jay , Eugénie Terreaux , Jean-Philippe Ovarlez , Frédéric Pascal

Providing efficient and accurate parametrizations for model reduction is a key goal in many areas of science and technology. Here we present a strong link between data-driven and theoretical approaches to achieving this goal. Formal…

Chaotic Dynamics · Physics 2021-06-02 Manuel Santos Gutiérrez , Valerio Lucarini , Mickaël D. Chekroun , Michael Ghil

We study the feature-based newsvendor problem, in which a decision-maker has access to historical data consisting of demand observations and exogenous features. In this setting, we investigate feature selection, aiming to derive sparse,…

Machine Learning · Computer Science 2022-09-13 Breno Serrano , Stefan Minner , Maximilian Schiffer , Thibaut Vidal

There is a great need for robust techniques in data mining and machine learning contexts where many standard techniques such as principal component analysis and linear discriminant analysis are inherently susceptible to outliers.…

Methodology · Statistics 2015-09-28 Garth Tarr , Samuel Müller , Neville C. Weber

Competing risk analysis considers event times due to multiple causes, or of more than one event types. Commonly used regression models for such data include 1) cause-specific hazards model, which focuses on modeling one type of event while…

Applications · Statistics 2017-04-27 Jiayi Hou , Anthony Paravati , Ronghui Xu , James Murphy

Causal effect estimation from observational data is a crucial but challenging task. Currently, only a limited number of data-driven causal effect estimation methods are available. These methods either provide only a bound estimation of the…

Methodology · Statistics 2020-11-10 Debo Cheng , Jiuyong Li , Lin Liu , Kui Yu , Thuc Duy Lee , Jixue Liu

Current causal discovery approaches require restrictive model assumptions in the absence of interventional data to ensure structure identifiability. These assumptions often do not hold in real-world applications leading to a loss of…

Machine Learning · Statistics 2025-06-25 Anish Dhir , Ruby Sedgwick , Avinash Kori , Ben Glocker , Mark van der Wilk

An efficient monotone data augmentation (MDA) algorithm is proposed for missing data imputation for incomplete multivariate nonnormal data that may contain variables of different types, and are modeled by a sequence of regression models…

Methodology · Statistics 2018-11-21 Yongqiang Tang

In health-pollution cohort studies, accurate predictions of pollutant concentrations at new locations are needed, since the locations of fixed monitoring sites and study participants are often spatially misaligned. For multi-pollution data,…

Applications · Statistics 2022-01-24 Phuong T. Vu , Adam A. Szpiro , Noah Simon

Matrix completion problem has been investigated under many different conditions since Netflix announced the Netflix Prize problem. Many research work has been done in the field once it has been discovered that many real life dataset could…

Machine Learning · Computer Science 2022-04-06 Jafar Jafarov

In panel data subject to nonignorable attrition, auxiliary (refreshment) sampling may restore full identification under weak assumptions on the attrition process. Despite their generality, these identification strategies have seen limited…

Econometrics · Economics 2025-12-16 Grigory Franguridi , Jinyong Hahn , Pierre Hoonhout , Arie Kapteyn , Geert Ridder