English
Related papers

Related papers: Regularized Targeted Maximum Likelihood Estimation…

200 papers

The Highly Adaptive Lasso (HAL) is a nonparametric regression method that achieves almost dimension-free convergence rates under minimal smoothness assumptions, but its implementation can be computationally prohibitive in high dimensions…

Machine Learning · Statistics 2026-05-06 Mingxun Wang , Alejandro Schuler , Mark van der Laan , Carlos García Meixide

Modern deep neural networks are powerful predictive tools yet often lack valid inference for causal parameters, such as treatment effects or entire survival curves. While frameworks like Double Machine Learning (DML) and Targeted Maximum…

Machine Learning · Computer Science 2025-07-17 Yi Li , David Mccoy , Nolan Gunter , Kaitlyn Lee , Alejandro Schuler , Mark van der Laan

Mixtures-of-Experts models and their maximum likelihood estimation (MLE) via the EM algorithm have been thoroughly studied in the statistics and machine learning literature. They are subject of a growing investigation in the context of…

Machine Learning · Statistics 2019-09-13 Faïcel Chamroukhi , Florian Lecocq , Hien D. Nguyen

This paper develops a unified estimation framework, the Maximum Ideal Likelihood Estimation (MILE), for general parametric models with latent variables. Unlike traditional approaches relying on the marginal likelihood of the observed data,…

Statistics Theory · Mathematics 2025-10-08 Yizhou Cai , Ting Fung Ma

We present an efficient algorithm for maximum likelihood estimation (MLE) of exponential family models, with a general parametrization of the energy function that includes neural networks. We exploit the primal-dual view of the MLE with a…

Machine Learning · Computer Science 2020-04-01 Bo Dai , Zhen Liu , Hanjun Dai , Niao He , Arthur Gretton , Le Song , Dale Schuurmans

Structured Latent Attribute Models (SLAMs) are a family of discrete latent variable models widely used in education, psychology, and epidemiology to model multivariate categorical data. A SLAM assumes that multiple discrete latent…

Methodology · Statistics 2021-07-12 Yuqi Gu , Gongjun Xu

The widespread use of generative models has created a feedback loop, in which each generation of models is trained on data partially produced by its predecessors. This process has raised concerns about model collapse: A critical degradation…

Machine Learning · Statistics 2026-03-27 Daniel Barzilai , Ohad Shamir

This paper considers one-step targeted maximum likelihood estimation method for general competing risks and survival analysis settings where event times take place on the positive real line R+ and are subject to right-censoring. Our…

Methodology · Statistics 2021-09-02 Helene C. W. Rytgaard , Mark J. van der Laan

As Large Language Models (LLMs) become increasingly integrated into real-world, autonomous applications, relying on static, pre-annotated references for evaluation poses significant challenges in cost, scalability, and completeness. We…

Computation and Language · Computer Science 2025-06-23 Sher Badshah , Ali Emami , Hassan Sajjad

We study nonparametric maximum likelihood estimation of probability densities under a total variation (TV) type penalty, sectional variation norm (also named as Hardy-Krause variation). TV regularization has a long history in regression and…

Statistics Theory · Mathematics 2026-02-19 Yilong Hou , Zhengpu Zhao , Yi Li , Mark van der Laan

Accumulated Local Effect (ALE) is a method for accurately estimating feature effects, overcoming fundamental failure modes of previously-existed methods, such as Partial Dependence Plots. However, ALE's approximation, i.e. the method for…

Machine Learning · Computer Science 2022-10-11 Vasilis Gkolemis , Theodore Dalamagas , Christos Diou

We provide a new interpretation of Hessian locally linear embedding (HLLE), revealing that it is essentially a variant way to implement the same idea of locally linear embedding (LLE). Based on the new interpretation, a substantial…

Machine Learning · Statistics 2021-12-17 Liren Lin , Chih-Wei Chen

We often seek to estimate the impact of an exposure naturally occurring or randomly assigned at the cluster-level. For example, the literature on neighborhood determinants of health continues to grow. Likewise, community randomized trials…

Methodology · Statistics 2021-07-08 Laura B. Balzer , Wenjing Zheng , Mark J. van der Laan , Maya L. Petersen

Exact MLE for generalized linear mixed models (GLMMs) is a long-standing problem unsolved until today. The proposed research solves the problem. In this problem, the main difficulty is caused by intractable integrals in the likelihood…

Methodology · Statistics 2024-10-14 Tonglin Zhang

Propensity score (PS) based estimators are increasingly used for causal inference in observational studies. However, model selection for PS estimation in high-dimensional data has received little attention. In these settings, PS models have…

This study focuses on the estimation of the Emax dose-response model, a widely utilized framework in clinical trials, agriculture, and environmental experiments. Existing challenges in obtaining maximum likelihood estimates (MLE) for model…

Methodology · Statistics 2025-06-11 Giacomo Aletti , Nancy Flournoy , Caterina May , Chiara Tommasi

We consider the problem of estimating the average treatment effect (ATE) when both randomized control trial (RCT) data and external real-world data (RWD) are available. We decompose the ATE estimand as the difference between a pooled-ATE…

Methodology · Statistics 2025-01-22 Mark van der Laan , Sky Qiu , Jens Magelund Tarp , Lars van der Laan

This work studies the properties of the maximum likelihood estimator (MLE) of a non-linear model with Gaussian errors and multidimensional parameter. The observations are collected in a two-stage experimental design and are dependent since…

Statistics Theory · Mathematics 2019-11-01 Nancy Flournoy , Caterina May , Chiara Tommasi

Accelerated life-testing (ALT) is a very useful technique for examining the reliability of highly reliable products. It allows testing the products at higher than usual stress conditions to induce failures more quickly and economically than…

Statistics Theory · Mathematics 2020-05-15 Aida Calviño

Auto-regressive sequence generative models trained by Maximum Likelihood Estimation suffer the exposure bias problem in practical finite sample scenarios. The crux is that the number of training samples for Maximum Likelihood Estimation is…

Machine Learning · Statistics 2020-07-14 Yuxuan Song , Ning Miao , Hao Zhou , Lantao Yu , Mingxuan Wang , Lei Li